Spendwaise
Agent context optimization

Agent context optimization: how to stop context bloat in multi-step agents

Agent context optimization is the practice of deciding, at every step of an agent loop, which tool outputs, memories and retrieved documents are worth sending to the model. Without it, context compounds across steps and most of the tokens the model reads never contribute to the task. The fix is to filter, summarize, deduplicate and budget context before each model call, then verify that task success holds.

Updated September 22, 202612 min readBy the Spendwaise team
Key takeaways
  • Agent context grows with every step because each model call re-reads the full history, including every previous tool result. A 10-step run can read the same payload ten times.
  • Most agent waste comes from four places: raw JSON tool payloads, MCP responses, stale intermediate state and repeated retrievals of the same content.
  • Six patterns fix it: filter tool output before it enters context, summarize completed sub-tasks, retrieve memory by task instead of replaying history, set a per-step token budget, deduplicate across steps and keep only task-relevant state.
  • Optimize context before inference, not after. Cheap relevance decisions on context cost far less than paying the generation model to read it.
  • Measure both sides: input tokens per step and task success rate on a fixed evaluation set. A reduction that lowers success is not a saving.

Why agent context grows with every step

Agent context grows because every model call in an agent loop re-reads the system prompt, the full conversation so far and every tool result that has been appended since the run started. Nothing is removed by default, so the prompt for step N contains the output of steps 1 through N minus 1.

A single chat request has a fixed prompt. An agent does not. Each iteration adds a model message, one or more tool calls and their results, and often a fresh retrieval. Because most agent frameworks keep the message list as the state, the entire list is serialized into the next request. The cost is quadratic in the number of steps: the total tokens read across a run is the sum of an increasing sequence, not a flat rate per step.

The arithmetic for a 10-step run

Take an illustrative agent with a 2,000-token system prompt and tools that return about 3,000 tokens per call. If nothing is filtered, the context at each step is the system prompt plus everything accumulated so far.

StepContext read by the modelCumulative tokens read
12,000 + 0 = 2,0002,000
22,000 + 3,000 = 5,0007,000
32,000 + 6,000 = 8,00015,000
52,000 + 12,000 = 14,00040,000
102,000 + 27,000 = 29,000155,000

Ten steps that produced 27,000 tokens of tool output cause the model to read 155,000 input tokens in total. The first tool result is read nine times. If the same run kept each step to a 6,000-token budget, the total would be at most 10 × 6,000 = 60,000 tokens, and the model would still see the parts of each result that mattered. These numbers are illustrative; the shape of the curve is not.

Where agent context waste comes from

Agent context waste comes from information that enters the prompt in the shape it arrived in, rather than in the shape the task needs. Five sources account for most of it.

Raw JSON tool payloads

A CRM lookup returns 40 fields per record when the agent needs three. A search API returns ranking metadata, snippets, URLs and timestamps for 20 results when the agent needs two. Serialized JSON is also token-expensive: keys, quotes, braces and whitespace routinely double the token count of the useful values.

MCP responses

Model Context Protocol servers expose tools that were designed for completeness, not for prompt size. A file listing, a database query or an issue tracker search often returns hundreds of rows because the server cannot know what the agent will do with them. The agent framework appends the whole response to context.

Stale intermediate state

Once a sub-task is finished, the reasoning that produced it is no longer useful. The plan the agent drafted at step 1, the three failed attempts at step 4 and the raw data behind a number it has already computed all stay in the message list and are re-read on every later step.

Repeated retrievals

Multi-step agents retrieve more than once. Each retrieval for a related query returns overlapping chunks, and the overlap is appended again. The same product page, policy paragraph or log excerpt ends up in context three or four times with slightly different boundaries.

Long chat history

In assistants that run over a long session, turns from an hour ago are still present in every request. Most of them concern tasks that are complete. Anthropic's guidance on effective context engineering treats context as a finite resource with diminishing returns, and history is the source that most often exceeds its useful size.

Source, typical waste and the fix

Each context source in an agent has a characteristic form of waste and a matching fix. The table below is the map most teams end up drawing after their first cost review.

Context sourceTypical wasteFix
Tool output (APIs, SQL, CRM)Full records with dozens of unused fields, verbose JSON encodingProject to the fields the task needs; render as compact text before appending
MCP responsesHundreds of rows or files returned for completenessFilter or summarize records against the current task before they enter context
Intermediate statePlans, drafts and failed attempts from finished sub-tasksReplace completed sub-task transcripts with a short result summary
Retrieved documentsOverlapping chunks from related queries across stepsDeduplicate across the run; keep one representative per fact
Conversation memoryOld turns replayed in every requestRetrieve only the history relevant to the current step
Web searchWhole pages when a paragraph answers the questionSelect relevant passages, not pages

Every row is the same move: decide what is worth sending before inference, using something much cheaper than the generation model. The broader case for that layer is covered in LLM context optimization; the retrieval-specific version is in RAG context compression.

Six patterns that fix agent context bloat

Agent context bloat is fixed by inserting a decision step between the sources that produce context and the model call that consumes it. Six patterns cover almost every agent, and they compose.

1. Filter tool output before it enters context

Treat every tool result as optional context, not as a fact the model must see. Score each record, row or field against the current step's goal and keep the subset that can change the model's next action. A search result with a relevance score below the threshold is dropped before serialization. This is the single highest-value pattern because tool output is both the largest source and the easiest to judge.

2. Summarize completed sub-tasks

When a sub-task closes, replace its transcript with a compact result: what was asked, what was found, what was decided. The reasoning and raw data are stored outside the prompt and can be re-fetched if a later step needs them. A 4,000-token investigation becomes a 120-token note.

3. Retrieve memory by task, not by recency

Do not replay the full session. Store turns and facts in a memory store and retrieve the ones relevant to the current step, the same way you retrieve documents. Redis describes this separation of short-term working context from long-term memory in its overview of AI agent context. The prompt then contains the last few turns plus what the current task needs, instead of everything that ever happened.

4. Set a per-step token budget

Give each step a ceiling for input context and fill it by value per token, not by arrival order. A budget turns a variable, compounding cost into a predictable one and forces the selection step to exist. Whole items should be kept or dropped; truncating the tail of a message is the wrong tool. Token budgets explains how to size them per agent.

5. Deduplicate across steps

Before appending a new retrieval or tool result, compare it with what is already in context. Near-identical chunks collapse into one. Records already summarized are not re-added. Dedup is cheap with embeddings or hashing and is the main defense against repeated retrievals.

6. Keep only task-relevant state

Build the prompt for each step from a state object, not from an append-only log. The state holds the goal, the current plan, the open sub-tasks and the evidence still needed. Serialization happens from the state, so anything that is no longer relevant simply stops being included. This is the shift from message history as state to state as the source of the prompt.

An agent loop with a context optimization step

The change to an agent loop is one step: after collecting context and before calling the model, select what fits the budget and the task. Everything else stays the same.

state = { goal, plan: [], evidence: [], done: [] }

for step in range(MAX_STEPS):
    # 1. gather context from every source
    items = []
    items += memory.retrieve(query=state.goal, k=20)
    items += state.evidence
    items += last_tool_results                 # raw payloads, not yet filtered

    # 2. optimize context before inference
    context, report = optimize(
        query=current_subtask(state),
        items=items,
        budget_tokens=6_000,
        dedupe=True,
        quality_guard=True,                    # keep items the answer depends on
    )
    log(report.tokens_before, report.tokens_after, report.dropped)

    # 3. call the model with the optimized context
    action = llm.generate(system=SYSTEM, state=state.summary(), context=context)

    # 4. execute, then store results outside the prompt
    last_tool_results = run_tools(action)
    state.evidence += project_fields(last_tool_results)   # keep useful fields only
    if action.closes_subtask:
        state.done.append(summarize(action))               # replace transcript with a note

    if action.is_final:
        return action.answer
Pseudo-code. Context is scored and budgeted before the expensive model call. The optimize step can be embeddings, a small classifier, heuristics or a service such as the Spendwaise API.

Two details matter. First, the optimize call takes the current sub-task as the query, not the original goal, because relevance changes as the agent progresses. Second, tool results go into state.evidence after field projection, so the raw payload is never serialized into a later prompt. If you want a hosted version of the optimize step, the Spendwaise API returns the optimized context plus a report of what was removed and why.

What it does to latency and cost

Optimizing agent context lowers both cost and latency, because the model reads fewer input tokens on every step and the selection step is cheaper than the tokens it removes.

Cost follows input tokens directly. Using the illustrative 10-step run above, moving from 155,000 to at most 60,000 tokens read is a 61% reduction in input tokens for that run. At a hypothetical input price of $3 per million tokens, that is $0.465 versus $0.180 per run. Multiply by runs per day to see why agents dominate LLM bills faster than chat does. The full cost picture, including output tokens, caching and model choice, is in the LLM cost optimization guide.

Latency follows input tokens too. Time to first token grows with prompt length, and in a 10-step run that delay is paid ten times. Relevance decisions on context run in parallel on small models and typically complete well before the generation model would have finished reading the tokens they removed. The net effect is a faster run, not a slower one, provided the optimization step is genuinely cheap.

  • Fewer input tokens per step: the direct saving, compounding across steps.
  • Lower time to first token per step: the model starts generating sooner.
  • Fewer context-window overflows: long runs stop failing at step 12 because the history no longer fits.
  • Better answers on late steps: less noise in the prompt means the model attends to the evidence that matters.

How to evaluate without breaking task success

Evaluate agent context optimization by measuring task success on a fixed set of runs before and after, alongside input tokens per step. Cost is the objective and quality is the constraint, never the other way around.

  1. Build an evaluation set of 50 to 200 real tasks with a known correct outcome or a rubric. Include long runs; short runs hide the compounding.
  2. Record the baseline: success rate, steps per run, input tokens per step, total tokens per run, wall-clock time.
  3. Turn on optimization with a conservative budget and a quality guard that keeps any item the current sub-task explicitly references.
  4. Re-run the set. Compare success rate first. If it dropped, inspect the removal log for the failed runs and find which drop caused it.
  5. Tighten the budget in steps only while success holds. Stop at the point where the next step costs accuracy.
  6. Keep the evaluation running in production on a sample of traffic, so drift in tools or data shows up as a quality alert rather than a support ticket.

The removal log is the debugging tool. Every dropped item should carry a reason: duplicate, off-topic, stale, over budget. When a run fails, the question is whether a needed item was dropped, and the log answers it in seconds. The quality guard exists to make that failure rare, and the savings report puts tokens saved next to quality so nobody optimizes blind.

Where agent context optimization fits in your stack

Agent context optimization sits between your tools, memory and retrieval on one side and the model call on the other. It does not replace the agent framework, the vector database or the LLM gateway.

LayerWhat it doesRelationship to context optimization
Agent frameworkRuns the loop, dispatches tools, holds stateHosts the optimize step; provides the context and the current sub-task
Vector database and retrieverFinds matching documentsFeeds results; a wider top-k becomes safe because selection happens after
LLM gatewayRoutes, caches, tracks spendSees fewer input tokens; caching and routing still apply
ObservabilityTraces calls, shows cost per stepDisplays the before and after tokens and the removal reasons

The wider discipline this belongs to is context engineering: treating the prompt as an engineered artifact assembled from sources, rather than a log that happens to be sent. Agents are where that discipline pays back fastest, because they are where context compounds. Send fewer, better tokens at every step and the rest of the stack gets cheaper and faster for free.

Frequently asked questions

What is agent context optimization?

Agent context optimization is the practice of selecting, at each step of an agent loop, which tool outputs, memories and retrieved documents are sent to the model. It filters, deduplicates, summarizes and budgets context before inference so the model reads fewer, more relevant tokens without losing what the task needs.

Why do agents use so many more tokens than chat?

Because each step re-reads the full history, including every previous tool result. Tokens are paid for on every subsequent step, so total input tokens grow roughly with the square of the number of steps. A 10-step run with 3,000-token tool results reads about 155,000 tokens if nothing is filtered.

Is truncating the message history enough?

No. Truncation drops the oldest content regardless of relevance, which often removes the goal or a key fact while keeping recent noise. Selecting by relevance to the current sub-task, with whole items kept or dropped, preserves what matters and removes what does not.

Does filtering tool output slow the agent down?

Usually the opposite. Relevance decisions run on small models in parallel and finish faster than the generation model would take to read the tokens they remove. Time to first token drops on every step, and the run finishes sooner.

How do I know optimization is not hurting answer quality?

Run a fixed evaluation set before and after, compare task success rate first and tokens second, and keep a removal log with a reason for every dropped item. Tighten the budget only while success holds. In production, sample traffic into the same evaluation so drift shows up as an alert.

Do I need a separate product to do this?

No. Field projection, summaries of completed sub-tasks, dedup and a per-step budget can all be built into your loop. A service such as Spendwaise packages the selection, budgeting, quality guard and reporting behind one call if you would rather not maintain it, but the patterns work either way.

Keep reading

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.