Spendwaise
Context optimization for production AI

Stop paying LLMs to read irrelevant context.

Spendwaise cleans, ranks and trims the context your AI sends to expensive models, in milliseconds, so you pay for fewer tokens without blindly cutting what the answer needs. One API call for RAG, one SDK line for agents.

See it live
Works with RAG, Agents, Tool calls, Web search and Chat history.
Overview
24h3d7d
Requests optimized
182,421
Input tokens sent
640M
was 1.82B
Context reduction
64.8%
Net cost avoided
$4,280
after scoring
Input tokens over timetokens insent to LLM
6:009:0012:0015:0018:0021:00
Waste removed by source
RAG612M → 198M
Agent state488M → 121M
MCP / tools301M → 74M
Chat history244M → 112M
Agents
support-triage5.3K-71%
research-planner9.8K-77%
sql-analyst8.0K-68%
email-drafter3.1K-61%

Your model isn't expensive. Your context is bloated.

RAG returns extra chunks. Agents dump raw tool output into every step, then replay it on every turn after. Web search brings back whole pages. The model reads all of it, and you pay for it, again and again.

Sources
RAGWeb searchMCP toolsChat history
before
18,420 tokens
Spendwaise
-71.4% context
after
5,260 tokens
Any model
OpenAI
Claude
Gemini
Mistral
DeepSeek
Llama
Cost per request $0.142 → $0.041
Tokens never sent 13,160
Answer quality -0.7% · groundedness +2.1%
RelevanceRedundancyBudget

Retrieve broadly.
Send narrowly.

  1. 01
    Retrieve
    Keep your existing vector search, web search or tools.
  2. 02
    Optimize
    Spendwaise strips noise, scores relevance, removes redundancy and applies your token budget.
  3. 03
    Generate
    Send the smaller, higher-signal context to OpenAI, Anthropic, Gemini or your model.
See the pipeline
full context58 items · 18,420 tok
cleanupnoise out · 1 ms58 items
relevance~100 ms58 → 21
redundancydeduplicate21 → 14
token budgetfit 6,00014 → 9
quality checkafter response0 ms
optimized context9 items · 5,260 tok
LatencyPer-step timing

Decided in milliseconds, before the expensive call.

Deciding has to be far cheaper and faster than reading, or it is not worth doing. Small questions go to a small, fast model; the generation model only reads what survives, which also shortens its time to first token.

One request, step by step125 ms
auth + project4 ms
cleanup1 ms
relevance118 ms
selection2 ms
response sent
quality check0 ms
request log0 ms
every request tagged< 500 ms500 ms to 1 s> 1 s
~100 ms
Relevance decisions
Judged by a small decision model that answers typed questions about every context item in parallel, in about 100 ms for most queries.
1–20 ms
Free local cleanup
Passing tests, repeated log lines, empty JSON fields and HTML markup are removed before anything else runs, measured at 1 to 20 ms per tool result on our demo agent.
0 ms
Nothing after the answer
The quality check, the request log and usage stats run after the response is sent. Your request never waits for them.
Agent SDKVercel AI SDK

Agents replay every tool result. Shrink it once.

A test log read at step 3 of a 20-step run is billed on every step after it. On our demo agent, a 7,209-token test run reached the model as 156 tokens, every failure kept.

Withoutrun_tests read at step 3
step 37,209
step 47,209
step 57,209
step 67,209
step 77,209
step 87,209
tokens read, steps 3 to 843,254
With the SDKrun_tests read at step 3
step 3156
step 4156
step 5156
step 6156
step 7156
step 8156
tokens read, steps 3 to 8936

The Spendwaise SDK (contextwaise on npm) wraps your Vercel AI SDK model and tools in one line.

Shrinks at birth
Each tool result is cut the moment the tool returns, before the model reads it, and stays that size for the rest of the run.
Nothing lost
The full output stays in your process. The model calls expand with a question to get any omitted part back.
Cache-safe
A result's view is decided once and replayed byte for byte, so your provider's prompt cache keeps hitting.
Explore the agent SDK
agent/run.ts
import { generateText, isStepCount } from "ai";
import { openai } from "@ai-sdk/openai";
import { createContextwaise } from "contextwaise";
import { withContextwaise } from "contextwaise/ai-sdk";

const cw = createContextwaise({ agent: "code-review" });
const { model, tools } = withContextwaise(cw, {
  model: openai("gpt-6.1-sol"),
  tools: { runTests, readFile, searchLogs },
});

// Every tool result reaches the model already shrunk.
// The full output stays with you, one expand() call away.
const { text } = await generateText({
  model, tools, prompt, stopWhen: isStepCount(10),
});
RAGAgentsMCPWeb

It is bigger than RAG.

AI Agents

The SDK shrinks every tool result the moment the tool returns, before it is replayed on every later step.

RAG

Retrieve more results, then pay your expensive model for only the evidence that matters.

Tool output

Test runs, logs, API JSON and file listings lose their noise first: passing tests, repeated lines, empty fields.

Web Search

Pages become text, then only the passages that answer the question reach the model.

MCP Tools

Stop dumping entire API responses into your agent context.

Long Documents

Spend context budget on the sections that can actually answer the question.

REST APIAny language

One call between retrieval and inference.

Keep your stack. Send the query and the context you would have sent; get back the context to send and a report of tokens saved, cost avoided and time spent per step. Kept items come back in your order, so your prompt cache keeps hitting.

Read the API
app/answer.ts
// Before your LLM call: send the context you would have sent
const res = await fetch(`${SPENDWAISE_URL}/api/v1/optimize`, {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.SPENDWAISE_KEY}` },
  body: JSON.stringify({
    query,
    candidates: [...docs, ...toolOutputs], // { text, source }
    agent: "support-triage", // per-agent budgets and stats
    model: "claude-sonnet-5-5", // prices the savings
  }),
});
const { context, report } = await res.json();

const answer = await llm.generate({ query, context });
console.log(report.tokens.savedPct, report.latencyMs); // 71.4 180
Quality checkRemoval log

Optimize cost without optimizing away the answer.

Every removal is logged with its reason: duplicate, low relevance, over budget, or noise cleaned. After each response, a quality check asks whether what was kept was enough to answer, and tunes that agent's threshold for the next request. Your response never waits for it.

How the quality check works
What was removedrequest 8f3a…
duplicate of #3-412 tok
Refund requests are processed within 5 to 7 business days…
low relevance-288 tok
Our office hours are Monday to Friday, 9am to 6pm CET…
cleaned-7,053 tok
336 passing tests and 6 progress bars removed from the test run…
over budget-630 tok
The 2023 pricing page listed three plans: Starter, Team…

Questions

The short ones are here. Anything else, email us.

What is Spendwaise?

A context layer between your sources (RAG, tools, web, memory) and your LLM. It removes noise, scores each context item against the current task, drops duplicates and fits a token budget before the expensive model call. Use it as a REST API from any language, or as an SDK inside your agent.

How fast is it?

Relevance is judged by a small decision model in about 100 ms for most queries, with every context item scored in parallel. The local cleanup takes milliseconds, and the quality check and logging run after the response. Every request records its timing per step and is tagged green under 500 ms, yellow up to 1 s and red above.

Does it work with the Vercel AI SDK?

Yes. withContextwaise wraps your model and tools in one line, using the AI SDK's own toModelOutput hook: your app keeps the raw tool output, the model reads the short version. Adapters for LangChain, the OpenAI Agents SDK, the Claude Agent SDK and MCP are next.

Will it break my prompt cache?

No. Providers bill a repeated prompt start at a fraction of the price only while it is byte-identical, so Spendwaise never edits what the model already saw: each result is decided once and replayed exactly, kept items stay in your order, and with a conversation id earlier decisions are remembered.

Will it cut context my model actually needs?

Errors are always kept, results that fit the budget pass whole, and nothing is deleted: the full output can be expanded back. After each response, a quality check asks whether the kept context was enough to answer, and the agent's threshold is adjusted for its next requests.

Which models does it work with?

Any. The final generation still uses OpenAI, Anthropic, Gemini, an open model or your own. Spendwaise only decides what reaches it, and never sits in your model's billing path.

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.