Spendwaise
Context optimization

LLM context optimization for RAG, agents and tool calls.

Retrieval finds results. Models reason over them. Spendwaise sits in between and decides which results are worth paying inference tokens for.

See it live
full context58 items · 18,420 tok
cleanupnoise out · 1 ms58 items
relevance~100 ms58 → 21
redundancydeduplicate21 → 14
token budgetfit 6,00014 → 9
quality checkafter response0 ms
optimized context9 items · 5,260 tok

Retrieval finds results. Models reason. Nothing in between decides what is worth reading.

Retrieval systems are optimized to find results. LLMs are optimized to reason over information. In most stacks, everything the first one finds is handed straight to the second one, and the model is paid to read all of it.

Spendwaise connects the two as a cheap decision layer. It scores your context against the current task, removes low-value information, handles redundancy, and enforces a token budget before the expensive model call.

Today: context gets dumped into the model
  1. 1User query
  2. 2Vector search / Web / MCP / Memory
  3. 320 to 100 retrieved chunks
  4. 4Everything is passed to the LLM
  5. 5Expensive inference
  6. 6Answer
With Spendwaise
  1. 1User query
  2. 2Vector search / Web / MCP / Memory
  3. 320 to 100 context items
  4. 4Spendwaise
  5. 55 to 15 high-signal items
  6. 6Expensive LLM
  7. 7Answer
Inside the layer

Six decisions, in order, before any expensive token is spent.

01
Free cleanup

Test output, logs, JSON, HTML and file listings lose their noise first: passing tests, repeated lines, empty fields, markup. Plain text is never touched.

02
Relevance decisions

A decision model answers one question per item, all in parallel: how much does this help answer the current question? About 100 ms for most queries.

03
Duplicate detection

Near-identical chunks and repeated tool output collapse into one, even when reworded.

04
Token-budget selection

The strongest items fill your budget first. Whole items are kept or dropped, never sliced, and they stay in the order you sent them.

05
Quality check

After the response, a decision model judges whether the kept context was enough to answer, and the agent's threshold adjusts for its next request.

06
Memory

With a conversation id, every decision is remembered, so content sent again is decided the same way and your prompt cache keeps hitting.

Why it is cheap

Cheap relevance decisions before expensive inference.

The economics only work if deciding is far cheaper than reading. Spendwaise asks small questions with small models and reserves the generation model for the reasoning it is actually good at.

Cheap decisions

A large language model is overkill for repeatedly answering small questions like "is this chunk relevant?". A small decision model built for typed questions answers them for a fraction of the cost.

Parallel evaluation

Every context item is scored against the same query at the same time, so adding more items barely adds latency: about 100 ms for most queries.

Model-agnostic

The final generation still uses OpenAI, Anthropic, Gemini, an open model or your customer's own model. Spendwaise only decides what reaches it.

The economic loop

Cheap selection happens before expensive reasoning.

Before
100k context tokens
→ 100k LLM input tokens

The application pays the expensive model to process context that may never contribute to the answer.

After
100k context tokens
→ 12k useful tokens → LLM

Selection costs a fraction of a cent per request. The exact reduction is benchmarked per workload, never assumed.

Every source

It is bigger than RAG.

Context bloat is not a retrieval problem. It shows up wherever information enters the prompt, and each source wastes tokens in its own way.

Context sourceTypical wasteWhat Spendwaise does
RAGTop-k contains weakly related chunksScore evidence against the current query and keep the useful subset
Web searchPages contain far more than the answer requiresSelect relevant passages instead of passing entire pages
AI agentsEvery step accumulates previous state and tool outputKeep only task-relevant state for the next model call
MCP / toolsAPIs return large JSON payloadsFilter or summarize the useful records before they enter context
Chat historyOld conversation turns remain in every requestRetrieve only history relevant to the current task
Long documentsLarge sections are retrieved when a few passages matterProgressively narrow pages to sections to chunks
Where it fits

Not a gateway. Not a reranker. Not a compressor.

Each of these solves a real problem. None of them answers the question Spendwaise is built for: which context is worth paying for on this request?

CategoryWhat it optimizesWhat Spendwaise adds
LLM gatewaysRoute, cache, track spend and observe calls.Spendwaise decides which context is worth paying for. Use both.
RerankersImprove the order of retrieval results.Spendwaise optimizes the whole context budget across RAG, web, tools and memory, not one retriever.
Prompt compressionShorten text with a compression model.Compression is one mechanism. Spendwaise sells the outcome: cost control with quality guardrails and measured savings.
Model routingSend each request to the cheapest capable model.Spendwaise sends fewer tokens to any model. Its own router (shadow mode) advises a cheaper model only when that pays back the prompt cache a switch loses.

Every context source, one optimization layer.

Relevance scoring

Each context item is scored against the current task by a small decision model that answers in about 100 ms, not by the expensive generation model.

Redundancy removal

Near-duplicate chunks, repeated tool output and restated history are collapsed before they reach the prompt.

Noise removed first

Test runs, logs, JSON, HTML and file listings lose passing tests, repeated lines, empty fields and markup before relevance is judged.

RAG context compression

Retrieve a wider top-k, then send only the subset that can answer the question.

Agent tool results

The agent SDK shrinks each tool result the moment it returns, before it is replayed on every later step.

Cache-safe

Kept items stay in your order and decisions never change once sent, so your prompt cache keeps hitting. See prompt caching.

Questions about context optimization

Does this replace my vector database or retriever?

No. Keep whatever finds your context today. Spendwaise runs after retrieval and before the model call, on the context you already have.

How is this different from setting a smaller top-k?

A smaller top-k throws away results before anyone has looked at them. Spendwaise lets you retrieve a wider set, then keeps only the items that actually help, which usually improves both cost and answer quality.

What happens when the model needs context that was removed?

Errors are always kept and every removal is logged with its reason. After each response, the quality check asks whether the kept context was enough; if not, that agent keeps more on its next requests. For agents, the full output also stays one expand call away.

Which decision model is used?

A small decision model that answers typed questions with calibrated probabilities instead of generating text. If it is unavailable, a local word-matching scorer takes over with a conservative bar.

How much does it save?

It depends on how bloated your context is. Start on the free plan and the dashboard reports your own before and after numbers rather than a marketing percentage.

Does it break prompt caching?

No. Kept items stay in your order and decisions never change once sent. See prompt caching.

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.