Spendwaise
LLM cost optimization

LLM cost optimization: a production playbook for cutting inference spend

LLM cost optimization means paying for fewer tokens per request without lowering answer quality. In most RAG and agent applications, input tokens make up the large majority of spend, so the highest-impact lever is sending less irrelevant context to the model. Measure per-request cost first, cut input context, then add caching, routing and output limits in that order.

Updated September 22, 202612 min readBy the Spendwaise team
Key takeaways
  • Input tokens usually dominate cost in RAG and agent apps because every request re-sends retrieved chunks, tool output and history while the answer stays short.
  • Measure before changing anything: cost per request, input and output tokens per request, and the share of context that actually influenced the answer.
  • The levers, in order of impact for context-heavy apps: cut input context, cache stable prefixes, route by difficulty, cap output, batch offline work, clean prompts.
  • Every reduction needs a quality constraint. Track answer quality on an evaluation set alongside spend, and treat quality regressions as a failed change.
  • Reduce token spend without blindly reducing context: score and select what reaches the model instead of shrinking top-k or truncating history.

The LLM cost equation

The cost of one LLM request is input tokens times the input price, plus output tokens times the output price, with cached input tokens billed at a discount. Every optimization technique works by shrinking one of those four terms or by moving a request to a cheaper price list.

cost_per_request =
    uncached_input_tokens * input_price
  + cached_input_tokens   * cached_input_price
  + output_tokens         * output_price

monthly_cost = sum(cost_per_request) over all requests
The full cost equation. Model tier changes the three prices; everything else changes the token counts.

Output tokens are priced several times higher than input tokens on most frontier models, which leads many teams to focus on shorter answers. That instinct is right for chat products with short prompts. It is wrong for RAG, agent and tool-using applications, where a typical request carries 10,000 to 50,000 input tokens and produces 200 to 800 output tokens. At that ratio, input spend exceeds output spend even after the price difference.

A worked example

Take an illustrative support agent with 18,000 input tokens and 400 output tokens per request, priced at $3 per million input tokens and $15 per million output tokens. Input cost is 18,000 × $3 / 1,000,000 = $0.054. Output cost is 400 × $15 / 1,000,000 = $0.006. Input is 90% of the bill. Halving output length saves $0.003 per request. Removing 60% of the input context saves $0.032 per request, more than ten times as much.

Why input tokens dominate in RAG and agent apps

Input tokens dominate because retrieval and orchestration layers are built to be generous, and the model is charged for everything they hand over. Each source of context adds tokens that were never checked for relevance to the current request.

Context sourceHow tokens pile upTypical share of input (illustrative)
Retrieved chunks (RAG)Top-k returns weakly related chunks alongside the useful ones30 to 60%
Tool and API outputFull JSON payloads are pasted into the prompt as-is10 to 40%
Conversation historyEvery prior turn is re-sent on every request10 to 30%
System prompt and instructionsGrows with every edge case a team patches5 to 15%
Web search resultsWhole pages are included when a few passages matter0 to 30%

The shares above are illustrative ranges, and your own distribution will differ. The point is structural: none of these sources decides what is worth reading. A vector database returns the nearest k vectors. A tool returns its whole response. A chat framework appends the whole transcript. The generation model is the first component that evaluates relevance, and by then the tokens are already paid for.

Anthropic's guidance on effective context engineering frames context as a finite resource with diminishing returns: past a point, more tokens do not improve answers and can degrade them. That makes context reduction a quality lever as well as a cost lever, provided the reduction removes noise rather than evidence.

Step one: measure cost per request

Before optimizing anything, instrument every LLM call to record input tokens, cached tokens, output tokens, model and cost, then aggregate by feature, agent and customer. Without per-request numbers, teams optimize the lever they understand rather than the lever that matters.

  1. Log token counts from the provider response on every call. Do not estimate from string length.
  2. Tag each call with the feature, agent step and customer that triggered it.
  3. Break input tokens down by source: system prompt, retrieved context, tool output, history, user message.
  4. Compute cost per request and cost per successful outcome (a resolved ticket, a completed task), not just cost per token.
  5. Sample requests and ask a simple question of each context item: did this item influence the final answer? The share that did not is your addressable waste.

The last step is the one most teams skip. It is also the one that reveals whether the problem is volume, model choice or context bloat. If 70% of retrieved chunks never influence the answer, no amount of prompt tightening will beat fixing the retrieval-to-prompt boundary.

Lever one: cut input context before inference

The single highest-impact lever for context-heavy applications is to optimize context before inference: score every context item against the current request, drop what is irrelevant or redundant, and fit the rest into a token budget. This is different from retrieving less, which throws away results before anyone has looked at them.

Retrieve broadly, send narrowly

Lowering top-k is the naive version of this lever. It cuts tokens, but it also cuts recall, because the useful chunk is sometimes ranked ninth. The better pattern is to retrieve a wide set of results, then run a cheap relevance decision per item and send only the survivors. Cheap here means a lightweight classifier, cross-encoder or embedding score, not another call to the generation model.

  • Relevance filtering: keep an item only if it plausibly helps answer this request, not just if it is semantically near the query.
  • Deduplication: collapse near-duplicate chunks, repeated tool results and restated history into one representative.
  • Prioritization: order survivors by expected contribution so the strongest evidence appears first.
  • Token budgets: pack items by value per token into a fixed budget per agent or request. See token budgets for how per-agent limits keep cost predictable.
  • Quality guardrails: if removal would starve the answer, keep the item and flag the request. See quality guard.

Worked example

Take the same illustrative agent: 18,000 input tokens, of which 12,000 are retrieved chunks, 3,500 are tool output, 1,500 are history and 1,000 are the system prompt. Suppose relevance filtering keeps 4 of 12 chunks (4,000 tokens), tool output is reduced to the relevant records (900 tokens), and history is trimmed to the turns that matter (500 tokens). New input is 1,000 + 4,000 + 900 + 500 = 6,400 tokens, a 64% reduction. At $3 per million input tokens, cost per request drops from $0.054 to $0.019 on the input side. Across 500,000 requests a month that is $27,000 down to $9,600, before any other lever.

These numbers are arithmetic on assumed inputs, not a benchmark. The actual reduction depends on how bloated your context is today, which is exactly what the measurement step tells you. Tools like Spendwaise implement this layer as a single call between retrieval and inference and report the before and after tokens per request, but the technique works with any relevance scorer you can run cheaply.

For source-specific detail, see RAG context compression for retrieval pipelines and agent context optimization for multi-step agents where tool output and state accumulate.

Lever two: cache stable prefixes

Prompt caching lets the provider reuse the processed form of a prompt prefix across requests, billing the cached portion at a fraction of the normal input price. It pays off when a large block of tokens is identical across many calls, which is true of system prompts, tool definitions and long reference documents. Both Anthropic and OpenAI document their caching rules and discounts.

Caching has two constraints that limit how far it goes. First, only the stable prefix can be cached, so the prompt has to be ordered with static content first and per-request content last. Second, retrieved context, tool output and the user message change on every request, and those are usually the largest part of a RAG or agent prompt. Caching removes the cost of the parts that were already cheap to reason about, and leaves the variable context untouched.

Cutting context and caching are complementary. Cut the variable context first, then cache the stable prefix. Applying caching to a bloated prompt locks in the bloat at a discount.

Levers three to six: routing, output limits, batching, prompt hygiene

The remaining levers each save a meaningful but smaller share, and each carries a distinct quality risk. Apply them after context is under control, because a smaller model reading 18,000 tokens of noise is still an expensive, unreliable model.

Model routing

Routing sends easy requests to a cheaper model and hard ones to a frontier model. The saving equals the price gap times the share of traffic that can safely be routed down. The risk is misclassification: an easy-looking request that needed the stronger model produces a wrong answer that costs more than the tokens saved. Routing works best when the router has a clear signal, such as request type, and when the cheaper model is evaluated on that slice specifically.

Output limits

Setting a maximum output length and instructing the model to answer concisely reduces the highest-priced token type. Savings are bounded by how long answers are today. For agents, the larger output saving usually comes from reducing the number of steps, since each step is another full request with the entire context re-sent.

Batching

Batch APIs process requests asynchronously at a discounted price. This applies only to work that can wait, such as backfills, classification jobs and nightly summaries. Interactive traffic cannot use it.

Prompt hygiene

System prompts accumulate instructions for edge cases and rarely lose any. Reviewing them quarterly, removing dead instructions and moving rarely used reference material into retrieval typically trims a few hundred to a few thousand tokens per request. The saving is small per request but applies to every request and improves instruction following.

Comparing the levers

For applications where input tokens dominate, cutting context offers the largest saving for moderate effort, while routing and caching offer meaningful savings with narrower applicability. The table below summarizes typical ranges, which are illustrative and depend heavily on the starting point.

LeverEffortTypical impact on spend (illustrative)Quality riskApplies to
Cut input contextMediumLarge where context is bloated; often the biggest single leverLow with guardrails, high if done by truncationRAG, agents, tools, long chats
Prompt cachingLowModerate; limited to the stable prefixNoneLong system prompts, static references
Model routingMediumModerate; depends on the share of easy trafficMedium, from misrouted hard requestsMixed-difficulty traffic
Output limitsLowSmall for RAG, larger for verbose chatLow to medium, from truncated answersAll
BatchingLowDiscounted price on eligible work onlyNoneOffline and async jobs
Prompt hygieneLowSmall per request, applies everywhereLow, often improves qualityAll

Read the table as an ordering, not a forecast. The measurement step tells you which row is large for your workload. A chat product with short prompts and long answers should start with output limits and routing. A RAG or agent product should start with context.

A 30-day LLM cost optimization plan

A realistic plan spends the first week measuring, the second and third weeks on the largest lever, and the last week locking in savings with an evaluation gate. Each step is reversible and produces a number you can report.

  1. Days 1 to 5: instrument. Log tokens and cost per call, tagged by feature, agent and customer. Break input down by source. Build a baseline dashboard.
  2. Days 6 to 8: build an evaluation set. Collect 100 to 300 real requests with graded answers. This is the quality constraint every later change must pass.
  3. Days 9 to 12: sample for waste. For 50 requests, mark which context items influenced the answer. Quantify addressable input tokens by source.
  4. Days 13 to 20: cut context. Add a relevance filter and token budget between retrieval and inference on the highest-volume feature. Compare tokens, cost and eval scores against the baseline. Ship if quality holds.
  5. Days 21 to 24: cache and reorder. Move static content to the prompt prefix and enable caching. Confirm the cached share in provider responses.
  6. Days 25 to 28: route and cap. Pilot a cheaper model on one well-defined request type with its own eval slice. Set output limits where answers run long.
  7. Days 29 to 30: report. Publish before and after cost per request, tokens per request, and eval scores. Set alerts on both cost and quality so regressions surface early.

Teams that skip the evaluation set tend to ship the cost saving and discover the quality regression from customers. Building the set on days 6 to 8 is what makes the rest of the plan safe to execute quickly.

Measuring savings with a quality constraint

A cost optimization is only real if answer quality on a fixed evaluation set stays within an agreed tolerance. Report savings and quality together, per change, so nobody can claim one without the other.

The minimum report has four numbers per feature: input tokens per request before and after, cost per request before and after, an answer-quality score on the evaluation set before and after, and a groundedness or faithfulness score before and after. Groundedness matters because removing noise often improves it, which is evidence that the reduction removed the right tokens. A savings report that shows removed tokens alongside eval deltas makes this routine rather than a one-off analysis.

for change in [context_filter, caching, routing]:
    before = run_eval(baseline, eval_set)
    after  = run_eval(with(change), eval_set)
    assert after.quality >= before.quality - tolerance
    report(change, tokens_saved, cost_saved, after.quality - before.quality)
Treat every cost change as a gated experiment. The tolerance is a product decision, not a default.

Set the tolerance deliberately. A small drop may be acceptable for an internal tool and unacceptable for a customer-facing legal assistant. The point of optimizing for cost subject to a quality constraint is that the constraint is explicit, agreed and enforced. For the broader discipline of deciding what reaches the model, see context engineering.

Frequently asked questions

What is the biggest lever for LLM cost optimization?

For RAG, agent and tool-using applications, the biggest lever is reducing input context: filtering irrelevant and redundant retrieved chunks, tool output and history before the model call. Input tokens usually make up most of the bill in these apps, so cutting them by half or more outweighs shorter answers or a cheaper model. For short-prompt chat products, output limits and routing matter more.

Does prompt caching solve LLM costs?

Partly. Caching discounts the stable prefix of a prompt, such as the system prompt and tool definitions. It does not reduce the cost of retrieved context, tool output or history, which change on every request and are usually the largest part of a RAG or agent prompt. Cut the variable context first, then cache the stable prefix.

Is it cheaper to just use a smaller model?

Sometimes, for well-defined easy requests with their own evaluation. But a smaller model reading the same bloated context is still paying for every token and is more likely to be distracted by noise. Reducing context first makes routing safer and cheaper, because the cheaper model reads only relevant evidence.

How do I reduce LLM input tokens without hurting quality?

Score each context item for relevance to the current request, remove duplicates, and fit survivors into a token budget, rather than lowering top-k or truncating history. Then verify on a fixed evaluation set that answer quality and groundedness stay within tolerance. Removing noise often improves groundedness, which is the signal that the right tokens were removed.

How much can context optimization save?

It depends on how much of your current context influences answers. Teams that sample requests and mark which items mattered typically find a large share of retrieved and tool tokens contributed nothing, but the exact figure varies by workload. Measure your own before and after on real traffic rather than relying on a published percentage.

Should I optimize cost per token or cost per outcome?

Cost per outcome. A cheaper request that fails and triggers a retry, an escalation or a second agent step costs more than a slightly more expensive request that succeeds. Track cost per resolved ticket or completed task alongside cost per request, and let quality gate every change.

Keep reading

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.