Krenzo
← All posts
Engineering·6 min read

Your search API isn't the expensive part

By Krenzo Engineering

When a team first prices out a retrieval pipeline, they look at the per-call cost of search and stop there. It's the visible number. It's also usually under a tenth of what the pipeline actually costs to run.

Where the money goes

A single deep research step might issue eight searches. Each returns five pages. If those pages arrive as raw HTML — nav bars, cookie banners, footers, inline scripts — you are pushing somewhere north of 200,000 tokens into a model that charges by the token. The search calls cost cents. The context costs dollars.

The fix isn't a cheaper search provider. It's returning less, and returning the right less. Boilerplate removal alone typically cuts payload by half. Passage-level ranking — returning the three paragraphs that answer the query rather than the whole article — takes another chunk out.

Measure the ratio, not the price

The metric worth tracking is tokens-per-useful-answer. It captures the thing you actually care about and it makes the tradeoff legible: a search call that costs twice as much but returns a quarter of the tokens is a straightforward win, and the naive per-call comparison will never show you that.

This is why Krenzo cleans and ranks before returning rather than leaving it to you. It's not a convenience feature. It's the part that determines whether the economics work.