A feature ships. It works. Two months later someone opens the provider dashboard and the monthly bill has four digits more than anyone budgeted, and nobody can say which endpoint is responsible. The investigation usually ends in the same place: one code path sends the full document on every turn of a conversation, so a ten-turn session re-sends the same 30,000 tokens ten times. Nobody noticed because each individual call looked reasonable.
This is the defining property of LLM cost. It is not one expensive thing. It is a small per-call cost multiplied by a call volume nobody modelled, made worse by context that grows quadratically with conversation length. And the same structural facts that drive cost drive latency, which is why they are worth fixing together.
Do the arithmetic before you build
Most cost surprises are arithmetic that was never done. The calculation takes five minutes and it should be part of the design, not the retrospective.
Take a support assistant. Each request sends a 1,200-token system prompt, 6,000 tokens of retrieved context, a 300-token question, and generates 500 tokens. That is 7,500 input and 500 output per call. At 2,000 calls a day you are moving 15 million input tokens and 1 million output tokens daily — around 450 million input tokens a month.
The exact rate depends on your provider and model, but the ratios are stable and they are what matters. Output tokens typically cost several times more than input tokens. Frontier models cost roughly an order of magnitude more than small ones. Cached input, where supported, costs a fraction of uncached input. Those three ratios determine almost every optimisation decision you will make.
Run the same calculation at ten times the volume. If the answer is unacceptable, the architecture is wrong now, not later — you will not optimise a 10x cost problem away with prompt tweaks.
Context is the cost driver, not call count
Teams instinctively try to reduce the number of calls. Usually the bigger win is reducing what each call carries.
In the example above, retrieved context is 80% of input tokens. Sending twelve chunks instead of six halves nothing and doubles the dominant term. And as covered in our note on building RAG systems that work in production, more chunks frequently makes answers worse as well as more expensive — material in the middle of a long context is used less reliably than material at the edges. Retrieving fifty candidates and reranking down to six is cheaper and better than sending twenty unranked.
Conversation history is the other offender, and it is worse because it compounds. Naively appending every turn means turn ten re-sends turns one through nine. Total tokens across a session grow with the square of turn count. Fixes in ascending order of effort: cap history to the last few turns, summarise older turns into a compact running state, or — best — maintain structured state rather than a transcript, so turn ten sends a state object of a few hundred tokens instead of nine turns of prose.
Prompt caching is the highest-leverage change available
If your provider supports prompt caching and you are not using it, this is the first thing to fix. It typically requires no change to model, prompt content or output quality — only to the order in which you assemble the context.
Caching works on a shared prefix. The provider recognises that the beginning of your request is identical to a recent one and skips recomputing it, charging a reduced rate for the cached portion. The requirement is an exact-match prefix, which means everything stable must come first and everything variable must come last.
Most naive prompt assembly breaks this by putting a timestamp or a user identifier near the top. One variable token at position 40 invalidates the entire cacheable prefix behind it.
// Cache-hostile: the timestamp at the top invalidates everything after it.
const bad = [
'Request at ' + now + ' for user ' + userId, // variable, position 0
SYSTEM_PROMPT, // 1,200 tokens, stable
TOOL_DEFINITIONS, // 900 tokens, stable
retrievedContext,
question,
].join('\n\n');
// Cache-friendly: stable prefix first, longest-lived first,
// variable material strictly at the tail.
function assemble(retrievedContext: string, question: string, meta: RequestMeta) {
return [
SYSTEM_PROMPT, // stable across every request
TOOL_DEFINITIONS, // stable across every request
FEW_SHOT_EXAMPLES, // stable; changes only on deploy
retrievedContext, // varies per query
formatMeta(meta), // varies per request
question, // varies per request
].join('\n\n');
}
Order the stable material by how long it lives: things that change on deploy before things that change per user before things that change per request. In the worked example, the system prompt, tool definitions and examples might be 2,500 tokens of a 7,500-token request. Making that third of every call cacheable is a large recurring saving for an afternoon of work.
Stop using one model for everything
The single most common source of waste is routing every request to the most capable model because that was what the prototype used.
Real workloads are a mix. Classification, extraction from structured input, routing decisions, short rewrites and yes/no judgements are handled well by small models. Multi-step reasoning, ambiguous synthesis and difficult code generation need a large one. If 70% of your traffic is the first category and you serve all of it with a frontier model, you are paying roughly ten times more than necessary for the majority of your volume.
Two structures work in production:
- Static routing by task type. You know at the call site which kind of work this is. Configure the model per task rather than globally. This is unglamorous and captures most of the available saving.
- Escalation. Attempt with the small model, validate the output deterministically, escalate to the large model only on failure. Economical when the small model succeeds most of the time — if it succeeds 80% of the time, you pay 1.0 small calls plus 0.2 large calls instead of 1.0 large calls. If it succeeds 40% of the time, you are paying for both and you should just use the large model.
The trap in escalation is validation. It only works if you can check the cheap output with code — schema conformance, a test run, a parse, a constraint check. If the only way to tell whether the small model got it right is to ask the large model, you have built a more expensive system, not a cheaper one.
Latency is a different problem with overlapping fixes
Cost and latency share causes but not remedies, and they occasionally conflict. Worth separating.
Input tokens are processed in parallel; output tokens are generated one at a time. This asymmetry is the most useful thing to know about LLM latency. A request with 8,000 input tokens and 200 output tokens is usually faster than one with 1,000 input and 1,500 output. If a response feels slow, look at output length before you look at input size.
Which gives a concrete lever: constrain output. Ask for structured output rather than prose with explanation. Request the fields you need and nothing else. A prompt that says "respond with JSON matching this schema, no commentary" can cut output tokens by more than half, which reduces both latency and the more expensive half of your bill.
| Lever | Cost effect | Latency effect | Risk |
|---|---|---|---|
| Prompt caching | Large reduction on stable prefix | Improves time to first token | None if ordering is correct |
| Smaller model for simple tasks | Large | Large | Quality regression if misrouted |
| Fewer, reranked context chunks | Large | Moderate | Needs a reranker; usually improves quality |
| Structured, bounded output | Large — output is the costly side | Large | Minimal |
| Streaming | None | Perceived latency only | Complicates validation of the whole response |
| Semantic caching | Large where queries repeat | Large on hit | Serving a near-miss as an exact answer |
| Summarised history | Removes quadratic growth | Moderate | Loses detail; summarisation costs a call |
Streaming deserves its caveat. It changes nothing about cost or total completion time, but time to first token is what users experience as speed, and the difference between four seconds of blank screen and text appearing in 400ms is enormous perceptually. It conflicts with validating the complete response before display, so for anything where a malformed or unsafe answer matters, either validate incrementally or do not stream.
Semantic caching — reusing a previous answer for a sufficiently similar question — is powerful where query distribution is concentrated, which in support-style workloads it usually is. The danger is the similarity threshold. "How do I cancel my subscription" and "how do I cancel my order" are close in embedding space and have different answers. Set the threshold conservatively, scope cache keys by anything that changes the correct answer (tenant, locale, entitlement), and expire on content updates.
Measure per unit of work, not per month
A monthly total tells you that you have a problem. It never tells you where. The metric that drives decisions is cost per completed unit of work — per resolved ticket, per processed document, per answered question — attributed to the feature that caused it.
That means tagging every model call with the feature, the task type, the model and the prompt version, and recording input tokens, cached tokens, output tokens and latency. With that in place you can answer the questions that matter: which feature is 60% of spend, did last week's prompt change increase output length, what is our cache hit rate, what is the p95 latency of the escalation path.
Two guardrails belong in the same layer. A per-request token ceiling, because a pathological input should fail fast rather than send 300,000 tokens. And a per-tenant or per-user rate limit, because unbounded automated traffic against a metered API is a financial incident waiting to happen. Both are trivial to add on day one and awkward to retrofit after a bill arrives. A cost regression check in CI — flagging when a prompt change materially increases tokens per unit of work — catches the slow drift that nobody attributes to anything.
What to skip
Two things absorb effort and rarely pay in the way teams expect.
Self-hosting an open-weight model to avoid API costs looks compelling in a spreadsheet and often is not. You take on GPU capacity planning, batching, autoscaling for spiky traffic, evaluation, and version management — all of it engineering time. It becomes genuinely economical at sustained high volume, with steady load, and a task where a smaller open model is sufficient. At moderate or bursty volume, idle GPU capacity costs more than the API you replaced.
Aggressive prompt compression — stripping words to save input tokens — is usually poor value. Input is the cheap side, and prompts degrade in ways that are hard to detect without a solid evaluation set. Removing a genuinely redundant thousand-token section is fine. Rewriting instructions telegraphically to save fifty tokens risks quality for a rounding error.
What this means in practice
Instrument first. Until every call is tagged by feature and task type with token counts attached, every optimisation is a guess, and the distribution is almost never what the team predicts. A week of instrumentation routinely reveals that one endpoint nobody was worried about is most of the bill.
Then work in order of leverage: enable prompt caching and reorder your context to make it effective; route simple tasks to small models; constrain output with structured schemas; reduce context through reranking rather than through sending more. Add a per-request token ceiling and a rate limit before you need them. Only after all of that is it worth evaluating semantic caching or self-hosting.
Cost control is not a one-off exercise, because prompts change, context grows and usage patterns shift. Treat tokens per unit of work as a tracked metric with the same seriousness as p95 latency, and the problem stays boring. If you are looking at a bill that outgrew its forecast, we are glad to help work out where it is going — that diagnosis is routine in our AI engineering work, and the fix is usually structural rather than clever.