Playbooks

How to run a cost review before scaling inference

Before adding capacity or switching models, walk the request once and price every stage. Most cost reviews find the money in context size, retries and caching, not in the model choice.

The order to do it in

Cost reviews usually start with the model catalogue, because that is where the visible prices are. Start one step earlier instead: follow a single representative request from the user interface to the response and write down every stage that consumes paid resources. The model is one line among several, and rarely the largest controllable one.

Stage by stage

  • Retrieval. Embedding the query, searching the index, and any reranking pass. Reranking a large candidate set with a cross-encoder is frequently more expensive than the generation it improves, and the trade is rarely measured.
  • Context assembly. The retrieved chunks, the conversation history, the system prompt, the tool definitions, and the format instructions. Everything here is billed on every turn, whether or not it changed since the previous turn.
  • Generation. Input and output tokens, with output tokens typically priced higher and produced more slowly. Verbosity is a cost decision, not only a style preference.
  • Retries and timeouts. Duplicate work that does not appear in any usage chart and scales with the tail of the latency distribution.
  • Evaluation and monitoring. Continuous scoring runs, synthetic traffic, and the judge model, which is often the same size as the production model.

The four levers that move the number

Context discipline. In most retrieval systems this is the largest controllable cost. Deduplicate chunks, cap the number of passages, drop history that is no longer referenced, and put the stable part of the prompt first so a cache can be reused. Measure the average context size per request and put it on a dashboard; it drifts upward without anybody deciding to increase it.

Routing. Send the questions that do not need a large model to a small one. The routing decision itself must be cheap and measurable, and the fallback — when the small model is not confident — must be defined in advance rather than discovered in production.

Caching. Identical and near-identical questions are far more common than product teams expect, particularly in internal tools with a small user population. Cache at the level of the answer, with an explicit invalidation rule tied to the corpus.

Output length. Ask for the length you want. A model told to be thorough will be thorough on every request, and the output is billed at the higher rate.

Writing the model down

Record the inputs, not only the result: requests per period, average input and output tokens per request, retry ratio, cache hit rate, and the unit prices used. A cost model without its inputs is a guess wearing a spreadsheet, and it cannot be re-checked when one of the assumptions changes — which is exactly when the decision needs revisiting.

Then attribute cost to the units the business cares about. Cost per active user and cost per completed task both explain a bill; cost per token does not, because nobody decided to buy a token.

When to buy capacity instead

After the levers above, if the workload is genuinely large and steady, self-hosted capacity may win. The decision rests on utilisation: reserved hardware is cheap per hour and expensive per unused hour. Write down the utilisation you expect, the utilisation you can measure, and the threshold at which the comparison flips. Without the utilisation number, the comparison is an argument, not a calculation.

What to do about it

  • Price the whole request path, not the model line item
  • Context size is the largest controllable cost in most RAG systems
  • Write the assumptions down — a cost model without inputs is a guess