LLM cost optimization: a production playbook for cutting inference spend
LLM cost optimization means paying for fewer tokens per request without lowering answer quality. In most RAG and agent applications, input tokens make up the large majority of spend, so the highest-impact lever is sending less irrelevant context to the model. Measure per-request cost first, cut input context, then add caching, routing and output limits in that order.
- Input tokens usually dominate cost in RAG and agent apps because every request re-sends retrieved chunks, tool output and history while the answer stays short.
- Measure before changing anything: cost per request, input and output tokens per request, and the share of context that actually influenced the answer.
- The levers, in order of impact for context-heavy apps: cut input context, cache stable prefixes, route by difficulty, cap output, batch offline work, clean prompts.
- Every reduction needs a quality constraint. Track answer quality on an evaluation set alongside spend, and treat quality regressions as a failed change.
- Reduce token spend without blindly reducing context: score and select what reaches the model instead of shrinking top-k or truncating history.
The LLM cost equation
The cost of one LLM request is input tokens times the input price, plus output tokens times the output price, with cached input tokens billed at a discount. Every optimization technique works by shrinking one of those four terms or by moving a request to a cheaper price list.
cost_per_request =
uncached_input_tokens * input_price
+ cached_input_tokens * cached_input_price
+ output_tokens * output_price
monthly_cost = sum(cost_per_request) over all requestsOutput tokens are priced several times higher than input tokens on most frontier models, which leads many teams to focus on shorter answers. That instinct is right for chat products with short prompts. It is wrong for RAG, agent and tool-using applications, where a typical request carries 10,000 to 50,000 input tokens and produces 200 to 800 output tokens. At that ratio, input spend exceeds output spend even after the price difference.
A worked example
Take an illustrative support agent with 18,000 input tokens and 400 output tokens per request, priced at $3 per million input tokens and $15 per million output tokens. Input cost is 18,000 × $3 / 1,000,000 = $0.054. Output cost is 400 × $15 / 1,000,000 = $0.006. Input is 90% of the bill. Halving output length saves $0.003 per request. Removing 60% of the input context saves $0.032 per request, more than ten times as much.
Why input tokens dominate in RAG and agent apps
Input tokens dominate because retrieval and orchestration layers are built to be generous, and the model is charged for everything they hand over. Each source of context adds tokens that were never checked for relevance to the current request.
| Context source | How tokens pile up | Typical share of input (illustrative) |
|---|---|---|
| Retrieved chunks (RAG) | Top-k returns weakly related chunks alongside the useful ones | 30 to 60% |
| Tool and API output | Full JSON payloads are pasted into the prompt as-is | 10 to 40% |
| Conversation history | Every prior turn is re-sent on every request | 10 to 30% |
| System prompt and instructions | Grows with every edge case a team patches | 5 to 15% |
| Web search results | Whole pages are included when a few passages matter | 0 to 30% |
The shares above are illustrative ranges, and your own distribution will differ. The point is structural: none of these sources decides what is worth reading. A vector database returns the nearest k vectors. A tool returns its whole response. A chat framework appends the whole transcript. The generation model is the first component that evaluates relevance, and by then the tokens are already paid for.
Anthropic's guidance on effective context engineering frames context as a finite resource with diminishing returns: past a point, more tokens do not improve answers and can degrade them. That makes context reduction a quality lever as well as a cost lever, provided the reduction removes noise rather than evidence.
Step one: measure cost per request
Before optimizing anything, instrument every LLM call to record input tokens, cached tokens, output tokens, model and cost, then aggregate by feature, agent and customer. Without per-request numbers, teams optimize the lever they understand rather than the lever that matters.
- Log token counts from the provider response on every call. Do not estimate from string length.
- Tag each call with the feature, agent step and customer that triggered it.
- Break input tokens down by source: system prompt, retrieved context, tool output, history, user message.
- Compute cost per request and cost per successful outcome (a resolved ticket, a completed task), not just cost per token.
- Sample requests and ask a simple question of each context item: did this item influence the final answer? The share that did not is your addressable waste.
The last step is the one most teams skip. It is also the one that reveals whether the problem is volume, model choice or context bloat. If 70% of retrieved chunks never influence the answer, no amount of prompt tightening will beat fixing the retrieval-to-prompt boundary.
Lever one: cut input context before inference
The single highest-impact lever for context-heavy applications is to optimize context before inference: score every context item against the current request, drop what is irrelevant or redundant, and fit the rest into a token budget. This is different from retrieving less, which throws away results before anyone has looked at them.
Retrieve broadly, send narrowly
Lowering top-k is the naive version of this lever. It cuts tokens, but it also cuts recall, because the useful chunk is sometimes ranked ninth. The better pattern is to retrieve a wide set of results, then run a cheap relevance decision per item and send only the survivors. Cheap here means a lightweight classifier, cross-encoder or embedding score, not another call to the generation model.
- Relevance filtering: keep an item only if it plausibly helps answer this request, not just if it is semantically near the query.
- Deduplication: collapse near-duplicate chunks, repeated tool results and restated history into one representative.
- Prioritization: order survivors by expected contribution so the strongest evidence appears first.
- Token budgets: pack items by value per token into a fixed budget per agent or request. See token budgets for how per-agent limits keep cost predictable.
- Quality guardrails: if removal would starve the answer, keep the item and flag the request. See quality guard.
Worked example
Take the same illustrative agent: 18,000 input tokens, of which 12,000 are retrieved chunks, 3,500 are tool output, 1,500 are history and 1,000 are the system prompt. Suppose relevance filtering keeps 4 of 12 chunks (4,000 tokens), tool output is reduced to the relevant records (900 tokens), and history is trimmed to the turns that matter (500 tokens). New input is 1,000 + 4,000 + 900 + 500 = 6,400 tokens, a 64% reduction. At $3 per million input tokens, cost per request drops from $0.054 to $0.019 on the input side. Across 500,000 requests a month that is $27,000 down to $9,600, before any other lever.
These numbers are arithmetic on assumed inputs, not a benchmark. The actual reduction depends on how bloated your context is today, which is exactly what the measurement step tells you. Tools like Spendwaise implement this layer as a single call between retrieval and inference and report the before and after tokens per request, but the technique works with any relevance scorer you can run cheaply.
For source-specific detail, see RAG context compression for retrieval pipelines and agent context optimization for multi-step agents where tool output and state accumulate.
Lever two: cache stable prefixes
Prompt caching lets the provider reuse the processed form of a prompt prefix across requests, billing the cached portion at a fraction of the normal input price. It pays off when a large block of tokens is identical across many calls, which is true of system prompts, tool definitions and long reference documents. Both Anthropic and OpenAI document their caching rules and discounts.
Caching has two constraints that limit how far it goes. First, only the stable prefix can be cached, so the prompt has to be ordered with static content first and per-request content last. Second, retrieved context, tool output and the user message change on every request, and those are usually the largest part of a RAG or agent prompt. Caching removes the cost of the parts that were already cheap to reason about, and leaves the variable context untouched.
Cutting context and caching are complementary. Cut the variable context first, then cache the stable prefix. Applying caching to a bloated prompt locks in the bloat at a discount.
Levers three to six: routing, output limits, batching, prompt hygiene
The remaining levers each save a meaningful but smaller share, and each carries a distinct quality risk. Apply them after context is under control, because a smaller model reading 18,000 tokens of noise is still an expensive, unreliable model.
Model routing
Routing sends easy requests to a cheaper model and hard ones to a frontier model. The saving equals the price gap times the share of traffic that can safely be routed down. The risk is misclassification: an easy-looking request that needed the stronger model produces a wrong answer that costs more than the tokens saved. Routing works best when the router has a clear signal, such as request type, and when the cheaper model is evaluated on that slice specifically.
Output limits
Setting a maximum output length and instructing the model to answer concisely reduces the highest-priced token type. Savings are bounded by how long answers are today. For agents, the larger output saving usually comes from reducing the number of steps, since each step is another full request with the entire context re-sent.
Batching
Batch APIs process requests asynchronously at a discounted price. This applies only to work that can wait, such as backfills, classification jobs and nightly summaries. Interactive traffic cannot use it.
Prompt hygiene
System prompts accumulate instructions for edge cases and rarely lose any. Reviewing them quarterly, removing dead instructions and moving rarely used reference material into retrieval typically trims a few hundred to a few thousand tokens per request. The saving is small per request but applies to every request and improves instruction following.
Comparing the levers
For applications where input tokens dominate, cutting context offers the largest saving for moderate effort, while routing and caching offer meaningful savings with narrower applicability. The table below summarizes typical ranges, which are illustrative and depend heavily on the starting point.
| Lever | Effort | Typical impact on spend (illustrative) | Quality risk | Applies to |
|---|---|---|---|---|
| Cut input context | Medium | Large where context is bloated; often the biggest single lever | Low with guardrails, high if done by truncation | RAG, agents, tools, long chats |
| Prompt caching | Low | Moderate; limited to the stable prefix | None | Long system prompts, static references |
| Model routing | Medium | Moderate; depends on the share of easy traffic | Medium, from misrouted hard requests | Mixed-difficulty traffic |
| Output limits | Low | Small for RAG, larger for verbose chat | Low to medium, from truncated answers | All |
| Batching | Low | Discounted price on eligible work only | None | Offline and async jobs |
| Prompt hygiene | Low | Small per request, applies everywhere | Low, often improves quality | All |
Read the table as an ordering, not a forecast. The measurement step tells you which row is large for your workload. A chat product with short prompts and long answers should start with output limits and routing. A RAG or agent product should start with context.
A 30-day LLM cost optimization plan
A realistic plan spends the first week measuring, the second and third weeks on the largest lever, and the last week locking in savings with an evaluation gate. Each step is reversible and produces a number you can report.
- Days 1 to 5: instrument. Log tokens and cost per call, tagged by feature, agent and customer. Break input down by source. Build a baseline dashboard.
- Days 6 to 8: build an evaluation set. Collect 100 to 300 real requests with graded answers. This is the quality constraint every later change must pass.
- Days 9 to 12: sample for waste. For 50 requests, mark which context items influenced the answer. Quantify addressable input tokens by source.
- Days 13 to 20: cut context. Add a relevance filter and token budget between retrieval and inference on the highest-volume feature. Compare tokens, cost and eval scores against the baseline. Ship if quality holds.
- Days 21 to 24: cache and reorder. Move static content to the prompt prefix and enable caching. Confirm the cached share in provider responses.
- Days 25 to 28: route and cap. Pilot a cheaper model on one well-defined request type with its own eval slice. Set output limits where answers run long.
- Days 29 to 30: report. Publish before and after cost per request, tokens per request, and eval scores. Set alerts on both cost and quality so regressions surface early.
Teams that skip the evaluation set tend to ship the cost saving and discover the quality regression from customers. Building the set on days 6 to 8 is what makes the rest of the plan safe to execute quickly.
Measuring savings with a quality constraint
A cost optimization is only real if answer quality on a fixed evaluation set stays within an agreed tolerance. Report savings and quality together, per change, so nobody can claim one without the other.
The minimum report has four numbers per feature: input tokens per request before and after, cost per request before and after, an answer-quality score on the evaluation set before and after, and a groundedness or faithfulness score before and after. Groundedness matters because removing noise often improves it, which is evidence that the reduction removed the right tokens. A savings report that shows removed tokens alongside eval deltas makes this routine rather than a one-off analysis.
for change in [context_filter, caching, routing]:
before = run_eval(baseline, eval_set)
after = run_eval(with(change), eval_set)
assert after.quality >= before.quality - tolerance
report(change, tokens_saved, cost_saved, after.quality - before.quality)Set the tolerance deliberately. A small drop may be acceptable for an internal tool and unacceptable for a customer-facing legal assistant. The point of optimizing for cost subject to a quality constraint is that the constraint is explicit, agreed and enforced. For the broader discipline of deciding what reaches the model, see context engineering.
Frequently asked questions
What is the biggest lever for LLM cost optimization?
For RAG, agent and tool-using applications, the biggest lever is reducing input context: filtering irrelevant and redundant retrieved chunks, tool output and history before the model call. Input tokens usually make up most of the bill in these apps, so cutting them by half or more outweighs shorter answers or a cheaper model. For short-prompt chat products, output limits and routing matter more.
Does prompt caching solve LLM costs?
Partly. Caching discounts the stable prefix of a prompt, such as the system prompt and tool definitions. It does not reduce the cost of retrieved context, tool output or history, which change on every request and are usually the largest part of a RAG or agent prompt. Cut the variable context first, then cache the stable prefix.
Is it cheaper to just use a smaller model?
Sometimes, for well-defined easy requests with their own evaluation. But a smaller model reading the same bloated context is still paying for every token and is more likely to be distracted by noise. Reducing context first makes routing safer and cheaper, because the cheaper model reads only relevant evidence.
How do I reduce LLM input tokens without hurting quality?
Score each context item for relevance to the current request, remove duplicates, and fit survivors into a token budget, rather than lowering top-k or truncating history. Then verify on a fixed evaluation set that answer quality and groundedness stay within tolerance. Removing noise often improves groundedness, which is the signal that the right tokens were removed.
How much can context optimization save?
It depends on how much of your current context influences answers. Teams that sample requests and mark which items mattered typically find a large share of retrieved and tool tokens contributed nothing, but the exact figure varies by workload. Measure your own before and after on real traffic rather than relying on a published percentage.
Should I optimize cost per token or cost per outcome?
Cost per outcome. A cheaper request that fails and triggers a retry, an escalation or a second agent step costs more than a slightly more expensive request that succeeds. Track cost per resolved ticket or completed task alongside cost per request, and let quality gate every change.