Pricing that costs less than the tokens it saves.
Start free with 10M context tokens a month. Paid plans are metered on the tokens you send, the same number your savings report measures against.
For evaluating on a real workload.
- 10M context tokens / month
- Agent SDK cleanup, always free
- Quality guard and savings report
- Per-agent budgets
- Community support
For production apps and agents.
- 100M context tokens / month
- $0.18 per extra 1M tokens
- Unlimited API keys
- Email support
For high-volume agents and platforms.
- 350M context tokens / month
- $0.15 per extra 1M tokens
- Everything in Pro
- Priority support
For platforms and regulated teams.
- Volume pricing
- Self-hosted deployment
- Custom decision models
- SLA and dedicated support
What you'd save, model by model.
Your model bills every input token. Spendwaise removes the ones it never needed, so you pay less, plan included.
| Model | Today | With Spendwaise | You save |
|---|---|---|---|
Claude Opus 5.5 $4 / 1M input | $523 | $186 $171 model + $15 Pro | $337 64% less |
Claude Sonnet 5.5 $2 / 1M input | $277 | $108 $93 model + $15 Pro | $169 61% less |
GPT-6.1 Sol $2 / 1M input | $212 | $85 $70 model + $15 Pro | $127 60% less |
Gemini 3.1 Pro $2 / 1M input | $227 | $92 $77 model + $15 Pro | $135 60% less |
DeepSeek V4 Pro $0.66 / 1M input | $68 | $37 $22 model + $15 Pro | $31 45% less |
GPT-6 Luna $0.1 / 1M input | $11 | $19 $4 model + $15 Pro | Costs more than it saves on a model this cheap |
You pay Spendwaise for 63.5M tokens a month: only the large tool outputs the SDK sends for scoring. Cleanup runs on your machine and is free. That fits Pro at $15 a month.
Monthly input cost at list prices, cache reads and writes included. Agent figures come from the run you can replay in the live demo: 5 tasks, 17 model calls on GPT-6.1 Sol. Output tokens are the same either way and left out. Your savings report measures your own numbers.
Every plan includes the quality guard, the savings report and per-agent budgets.
You pay for the context tokens you send, counted the way the dashboard counts them. A long tool output and a short chat turn cost what they weigh.
The agent SDK cleans every tool result on your machine, for free. Only large outputs go to the API for scoring, and only those count. In our recorded run that was a quarter of what the model would have read.
You keep calling your LLM provider directly. Spendwaise never sits in that billing path.
If a month's savings report shows less net LLM cost avoided than your plan price, we credit the difference. Applies to models priced from our catalog.
Teams that join during early access get 30% off Pro or Scale for their first year. Not combinable with yearly billing.
Pay yearly and save 20%: $144 a year for Pro, $468 for Scale.
Remove one call, or unwrap your agent, and your stack works exactly as before.
Questions about pricing
What counts as a context token?
The tokens in the query and context items you send to the API, counted at about four characters per token, the same count the dashboard shows. Tokens Spendwaise removes still count: deciding they can go is the work. With the agent SDK, only the tool outputs it sends for scoring count; resending a result on later steps is free.
What happens past my included tokens?
On Pro and Scale, extra tokens are billed per started million: $0.18 on Pro, $0.15 on Scale. Past about 290M tokens a month, Scale costs less than Pro with extra tokens. On Free, requests keep running and you are asked to upgrade; nothing is billed.
When does Pro pay for itself?
Illustrative: at the full Pro allowance of 100M tokens, if Spendwaise removes 70% of the context, 70M input tokens never reach your model. Pro pays for itself once your model costs more than about $0.22 per million input tokens. On a model at $2 per million, that month avoids about $140 of input spend.
Do you charge for LLM tokens?
No. You keep paying your model provider directly. Spendwaise decides what context reaches the model; it never resells or marks up model usage.
How do I know it is worth the price?
The savings report shows tokens removed, the LLM cost avoided and the quality impact on your own workload. Start on the free plan and compare the cost avoided with the plan price.
Do agent tool results count?
Only the ones the agent SDK sends to the API for relevance scoring, counted the same way. Results it handles locally (cleanup, small outputs) are free and never leave your machine.
Is the quality guard a paid add-on?
No. Every plan includes the quality guard and the savings report, because optimizing cost without measuring quality is not something we sell.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.