LLM context optimization for RAG, agents and tool calls.
Retrieval finds results. Models reason over them. Spendwaise sits in between and decides which results are worth paying inference tokens for.
Retrieval finds results. Models reason. Nothing in between decides what is worth reading.
Retrieval systems are optimized to find results. LLMs are optimized to reason over information. In most stacks, everything the first one finds is handed straight to the second one, and the model is paid to read all of it.
Spendwaise connects the two as a cheap decision layer. It scores your context against the current task, removes low-value information, handles redundancy, and enforces a token budget before the expensive model call.
- 1User query
- 2Vector search / Web / MCP / Memory
- 320 to 100 retrieved chunks
- 4Everything is passed to the LLM
- 5Expensive inference
- 6Answer
- 1User query
- 2Vector search / Web / MCP / Memory
- 320 to 100 context items
- 4Spendwaise
- 55 to 15 high-signal items
- 6Expensive LLM
- 7Answer
Six decisions, in order, before any expensive token is spent.
Test output, logs, JSON, HTML and file listings lose their noise first: passing tests, repeated lines, empty fields, markup. Plain text is never touched.
A decision model answers one question per item, all in parallel: how much does this help answer the current question? About 100 ms for most queries.
Near-identical chunks and repeated tool output collapse into one, even when reworded.
The strongest items fill your budget first. Whole items are kept or dropped, never sliced, and they stay in the order you sent them.
After the response, a decision model judges whether the kept context was enough to answer, and the agent's threshold adjusts for its next request.
With a conversation id, every decision is remembered, so content sent again is decided the same way and your prompt cache keeps hitting.
Cheap relevance decisions before expensive inference.
The economics only work if deciding is far cheaper than reading. Spendwaise asks small questions with small models and reserves the generation model for the reasoning it is actually good at.
A large language model is overkill for repeatedly answering small questions like "is this chunk relevant?". A small decision model built for typed questions answers them for a fraction of the cost.
Every context item is scored against the same query at the same time, so adding more items barely adds latency: about 100 ms for most queries.
The final generation still uses OpenAI, Anthropic, Gemini, an open model or your customer's own model. Spendwaise only decides what reaches it.
Cheap selection happens before expensive reasoning.
→ 100k LLM input tokens
The application pays the expensive model to process context that may never contribute to the answer.
→ 12k useful tokens → LLM
Selection costs a fraction of a cent per request. The exact reduction is benchmarked per workload, never assumed.
It is bigger than RAG.
Context bloat is not a retrieval problem. It shows up wherever information enters the prompt, and each source wastes tokens in its own way.
| Context source | Typical waste | What Spendwaise does |
|---|---|---|
| RAG | Top-k contains weakly related chunks | Score evidence against the current query and keep the useful subset |
| Web search | Pages contain far more than the answer requires | Select relevant passages instead of passing entire pages |
| AI agents | Every step accumulates previous state and tool output | Keep only task-relevant state for the next model call |
| MCP / tools | APIs return large JSON payloads | Filter or summarize the useful records before they enter context |
| Chat history | Old conversation turns remain in every request | Retrieve only history relevant to the current task |
| Long documents | Large sections are retrieved when a few passages matter | Progressively narrow pages to sections to chunks |
Not a gateway. Not a reranker. Not a compressor.
Each of these solves a real problem. None of them answers the question Spendwaise is built for: which context is worth paying for on this request?
| Category | What it optimizes | What Spendwaise adds |
|---|---|---|
| LLM gateways | Route, cache, track spend and observe calls. | Spendwaise decides which context is worth paying for. Use both. |
| Rerankers | Improve the order of retrieval results. | Spendwaise optimizes the whole context budget across RAG, web, tools and memory, not one retriever. |
| Prompt compression | Shorten text with a compression model. | Compression is one mechanism. Spendwaise sells the outcome: cost control with quality guardrails and measured savings. |
| Model routing | Send each request to the cheapest capable model. | Spendwaise sends fewer tokens to any model. Its own router (shadow mode) advises a cheaper model only when that pays back the prompt cache a switch loses. |
Every context source, one optimization layer.
Each context item is scored against the current task by a small decision model that answers in about 100 ms, not by the expensive generation model.
Near-duplicate chunks, repeated tool output and restated history are collapsed before they reach the prompt.
Test runs, logs, JSON, HTML and file listings lose passing tests, repeated lines, empty fields and markup before relevance is judged.
Retrieve a wider top-k, then send only the subset that can answer the question.
The agent SDK shrinks each tool result the moment it returns, before it is replayed on every later step.
Kept items stay in your order and decisions never change once sent, so your prompt cache keeps hitting. See prompt caching.
Questions about context optimization
Does this replace my vector database or retriever?
No. Keep whatever finds your context today. Spendwaise runs after retrieval and before the model call, on the context you already have.
How is this different from setting a smaller top-k?
A smaller top-k throws away results before anyone has looked at them. Spendwaise lets you retrieve a wider set, then keeps only the items that actually help, which usually improves both cost and answer quality.
What happens when the model needs context that was removed?
Errors are always kept and every removal is logged with its reason. After each response, the quality check asks whether the kept context was enough; if not, that agent keeps more on its next requests. For agents, the full output also stays one expand call away.
Which decision model is used?
A small decision model that answers typed questions with calibrated probabilities instead of generating text. If it is unavailable, a local word-matching scorer takes over with a conservative bar.
How much does it save?
It depends on how bloated your context is. Start on the free plan and the dashboard reports your own before and after numbers rather than a marketing percentage.
Does it break prompt caching?
No. Kept items stay in your order and decisions never change once sent. See prompt caching.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.