Stop paying LLMs to read irrelevant context.
Spendwaise cleans, ranks and trims the context your AI sends to expensive models, in milliseconds, so you pay for fewer tokens without blindly cutting what the answer needs. One API call for RAG, one SDK line for agents.
Your model isn't expensive. Your context is bloated.
RAG returns extra chunks. Agents dump raw tool output into every step, then replay it on every turn after. Web search brings back whole pages. The model reads all of it, and you pay for it, again and again.
Retrieve broadly.
Send narrowly.
- 01RetrieveKeep your existing vector search, web search or tools.
- 02OptimizeSpendwaise strips noise, scores relevance, removes redundancy and applies your token budget.
- 03GenerateSend the smaller, higher-signal context to OpenAI, Anthropic, Gemini or your model.
Decided in milliseconds, before the expensive call.
Deciding has to be far cheaper and faster than reading, or it is not worth doing. Small questions go to a small, fast model; the generation model only reads what survives, which also shortens its time to first token.
Agents replay every tool result. Shrink it once.
A test log read at step 3 of a 20-step run is billed on every step after it. On our demo agent, a 7,209-token test run reached the model as 156 tokens, every failure kept.
The Spendwaise SDK (contextwaise on npm) wraps your Vercel AI SDK model and tools in one line.
expand with a question to get any omitted part back.import { generateText, isStepCount } from "ai";
import { openai } from "@ai-sdk/openai";
import { createContextwaise } from "contextwaise";
import { withContextwaise } from "contextwaise/ai-sdk";
const cw = createContextwaise({ agent: "code-review" });
const { model, tools } = withContextwaise(cw, {
model: openai("gpt-6.1-sol"),
tools: { runTests, readFile, searchLogs },
});
// Every tool result reaches the model already shrunk.
// The full output stays with you, one expand() call away.
const { text } = await generateText({
model, tools, prompt, stopWhen: isStepCount(10),
});It is bigger than RAG.
The SDK shrinks every tool result the moment the tool returns, before it is replayed on every later step.
Retrieve more results, then pay your expensive model for only the evidence that matters.
Test runs, logs, API JSON and file listings lose their noise first: passing tests, repeated lines, empty fields.
Pages become text, then only the passages that answer the question reach the model.
Stop dumping entire API responses into your agent context.
Spend context budget on the sections that can actually answer the question.
One call between retrieval and inference.
Keep your stack. Send the query and the context you would have sent; get back the context to send and a report of tokens saved, cost avoided and time spent per step. Kept items come back in your order, so your prompt cache keeps hitting.
Read the API// Before your LLM call: send the context you would have sent
const res = await fetch(`${SPENDWAISE_URL}/api/v1/optimize`, {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SPENDWAISE_KEY}` },
body: JSON.stringify({
query,
candidates: [...docs, ...toolOutputs], // { text, source }
agent: "support-triage", // per-agent budgets and stats
model: "claude-sonnet-5-5", // prices the savings
}),
});
const { context, report } = await res.json();
const answer = await llm.generate({ query, context });
console.log(report.tokens.savedPct, report.latencyMs); // 71.4 180Optimize cost without optimizing away the answer.
Every removal is logged with its reason: duplicate, low relevance, over budget, or noise cleaned. After each response, a quality check asks whether what was kept was enough to answer, and tunes that agent's threshold for the next request. Your response never waits for it.
How the quality check worksQuestions
The short ones are here. Anything else, email us.
What is Spendwaise?
A context layer between your sources (RAG, tools, web, memory) and your LLM. It removes noise, scores each context item against the current task, drops duplicates and fits a token budget before the expensive model call. Use it as a REST API from any language, or as an SDK inside your agent.
How fast is it?
Relevance is judged by a small decision model in about 100 ms for most queries, with every context item scored in parallel. The local cleanup takes milliseconds, and the quality check and logging run after the response. Every request records its timing per step and is tagged green under 500 ms, yellow up to 1 s and red above.
Does it work with the Vercel AI SDK?
Yes. withContextwaise wraps your model and tools in one line, using the AI SDK's own toModelOutput hook: your app keeps the raw tool output, the model reads the short version. Adapters for LangChain, the OpenAI Agents SDK, the Claude Agent SDK and MCP are next.
Will it break my prompt cache?
No. Providers bill a repeated prompt start at a fraction of the price only while it is byte-identical, so Spendwaise never edits what the model already saw: each result is decided once and replayed exactly, kept items stay in your order, and with a conversation id earlier decisions are remembered.
Will it cut context my model actually needs?
Errors are always kept, results that fit the budget pass whole, and nothing is deleted: the full output can be expanded back. After each response, a quality check asks whether the kept context was enough to answer, and the agent's threshold is adjusted for its next requests.
Which models does it work with?
Any. The final generation still uses OpenAI, Anthropic, Gemini, an open model or your own. Spendwaise only decides what reaches it, and never sits in your model's billing path.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.