Measure LLM cost savings from context optimization.
The product measures its own value. Tokens saved, dollars avoided and quality preserved, per agent, per source, per month.
A saving you cannot measure is a marketing number.
Context optimization changes two things at once: what you pay and what the model knows. A report that shows only the first is half a result. The savings report puts tokens, dollars and quality in the same place, so the saving is something engineering can defend and finance can plan around.
It is also how Spendwaise measures its own value. If the report shows no saving on your workload, you should not be paying for it.
Every figure traces back to requests.
The baseline is the context you asked Spendwaise to optimize: every context item passed into the call. The result is the context it returned. Everything else is derived from those two numbers.
| Metric | How it is computed | What it tells you |
|---|---|---|
| Input tokens before | Tokens in all context items sent for optimization. | What the model would have read. |
| Input tokens after | Tokens in the optimized context returned. | What the model actually read. |
| Context reduction | One minus after divided by before. | How bloated the context was. |
| Cost avoided | Tokens saved, priced at each model's input rate. | The saving in dollars, before any caching discounts. |
| Quality check | A decision model's verdict, after the response, on whether the kept context was enough to answer. | Whether the saving cost you anything. |
| Latency | Time per step, tagged green under 500 ms, yellow to 1 s, red above. | What the optimization cost in time. |
Find out where the waste comes from.
Savings are broken down by context source, and the source that dominates tells you where to look next.
| If most waste comes from | It usually means | Read next |
|---|---|---|
| RAG | Top-k is wide and chunks overlap. | RAG context compression |
| Agent state and tool output | Each step replays everything before it. | Agent context optimization |
| Chat history | Old turns ride along on every request. | Context engineering |
| MCP tools and web search | Raw payloads and whole pages enter the prompt. | Context optimization |
One report for engineering and finance.
Engineering reads it per agent and per source to decide what to tune next. Finance reads it per month to attribute LLM spend and forecast it. Both read the same numbers, which is what stops the two conversations from drifting apart.
Every optimization response from the API already carries its own report: tokens and cost saved, the removal reasons and the timings, ready to log next to your own request logs.
A report your finance team can read.
Input tokens with and without optimization, aggregated across every request.
Token savings priced at each model's input rate, so the number is in dollars.
See where waste comes from: RAG, agent state, MCP tools, chat history or web search.
Attribute savings to the agents that generate them, with per-agent budgets and learned thresholds.
Each request shows the quality check's verdict next to its savings, so nobody optimizes blindly.
A dedicated view per tool: tokens returned, tokens the model read, and how often the agent asked for more.
Questions about savings report
Is cost avoided the same as my bill going down?
Not exactly. Cost avoided is measured against the full context you sent. If you already truncated context before Spendwaise, compare provider invoices before and after as well; the report shows what the optimizer removed, the invoice shows what you paid.
Does the report include answer quality?
Yes. Each request shows the quality check's verdict (was the kept context enough to answer?) next to its savings, and how it moved that agent's threshold.
Can I break savings down by agent and source?
Yes. Name the agent and label each context item with a source (rag, tools, web…) and every figure is split both ways.
Can I get the data programmatically?
Yes. Each optimization response includes its own report: tokens and cost before and after, per-item decisions, latency and timings.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.