Delete 62% of your agent's context. It gets better, not worse.
Three new studies agree: the tool history you are paying to keep is actively hurting the agent. Prune it, and cost and accuracy both move in the right direction.

Your agent remembers everything. Every tool call, every tool response, every half-formed plan, carried forward into every future step. It feels like the safe default: more context, better decisions. Three new papers put a number on that instinct, and the number should worry you.
On one enterprise benchmark, an agent that kept its full history completed 71.0% of tasks. An agent that threw away everything except its last five tool calls completed 79.0%. The agent that also kept a one-paragraph summary of what it threw away completed 91.6%, while using 62.7% fewer tokens and finishing in 60.2% less time. (Lodha et al., arXiv:2606.10209)
Two more studies, on different stacks and different domains, land in the same place. Xiao et al. (FSE 2026, arXiv:2509.23586) cut input tokens by 39.9 to 59.7% and total compute cost by 21.1 to 35.9% on a top-performing coding agent, by deleting useless, redundant, and expired trajectory information, with no loss in performance. Kang et al. (arXiv:2510.00615) cut peak token usage by 26 to 54% across AppWorld, OfficeBench, and multi-objective QA while improving success rates, and their compression scheme lifted smaller models by up to 46% on long-horizon tasks.
The pattern is consistent enough to state plainly: most of what your agent remembers is not helping it. Some of it is hurting it.
The findings
1. Throwing away history raised accuracy before anything clever was added
Lodha et al. studied GPT-5 driving Microsoft Dynamics 365 Finance and Operations through MCP tools on a 50-task hotel-expense benchmark, with results averaged over five independent runs. Their second configuration (C2) kept full conversation history: 71.0% complete itemization. Their third configuration (C3) kept only the last five tool call/response pairs and discarded the rest: 79.0%.
Nothing was added to compensate. The gain came from removal alone. The authors' explanation deserves attention: stale tool responses describe superseded form state, and the agent acts on them. Old context is not neutral. It is misinformation about a world that no longer exists.
2. Summarization bought the rest: 91.6% on 62.7% fewer tokens
The fourth configuration (C4) added an automated summary of the evicted pairs back into context. Complete itemization rose to 91.6%, with 99.64% of amounts correctly itemized. Against the full-history baseline, that is 20.6 percentage points of accuracy on 62.7% fewer tokens (1,480,996 down to 553,374) and 60.2% less wall-clock time (14.56 hours down to 5.79 hours per benchmark run).
The detail practitioners should not miss: summarization was not free. C4 consumed roughly 18,000 more tokens than C3 (553,374 vs 535,274) to buy its extra 12.6 points. Compression has a budget line of its own.
3. The effect is not one model's quirk
The paper reports confidence intervals and effect-size analysis, sensitivity over pruning and summary window sizes, and cross-model evidence with Claude Sonnet 4.5. The pruning-plus-summary advantage survived the model switch. That matters, because a finding that only holds on one provider's flagship is a demo, not a result.
4. On coding agents, 40 to 60% of input tokens are waste
Xiao et al. analyzed real agent trajectories and found useless, redundant, and expired information "widespread" across them. Their AgentDiet system removes that waste during execution on a top-performing coding agent. Across two LLMs and two benchmarks: input tokens down 39.9 to 59.7%, total computational cost down 21.1 to 35.9%, with agent performance maintained.
The wide range is itself informative. Waste is not a fixed fraction. It depends on the task, the model, and the agent's habits. If your agent retries aggressively or re-reads files it has already seen, you are probably at the high end.
5. Optimizing what to keep, from failures, beats static rules
Kang et al.'s ACON takes a different route. Instead of a fixed window, it iteratively refines compression guidelines in natural language from failure analysis of the agent, then distills the compressor into smaller models. On AppWorld, OfficeBench, and multi-objective QA: peak token usage down 26 to 54%, with success rates improving over existing compression baselines. Smaller language models gained up to 46% on long-horizon tasks, mostly by being distracted less.
Read findings 1 and 5 together and the message sharpens: the best context policy is not "keep more." It is "keep what the task still needs, and know what that is from where you have failed before."
What this means for data and AI practitioners
- Price your agent per successful task, not per call. Token-per-task numbers are the only ones that survive contact with production. Lodha's full-history agent was not 20 points worse in some abstract sense; it was 20 points worse while consuming 2.7 times the tokens.
- Treat stale tool output as a correctness bug, not a cost nuisance. Superseded form state, expired API responses, and already-fixed errors sitting in context are instructions the agent will follow. Pruning is error prevention, not just savings.
- Budget for the summarizer. Compression is not free. The best configuration in the Lodha study paid about 18K tokens for its summary. Measure the net, not the gross.
- If you run agents on small or cheap models, this matters more for you. ACON's 46% lift for smaller models suggests context discipline is a bigger lever than model upgrades for long-horizon work.
- Your eval harness should track tokens per task alongside success rate. One without the other hides the trade. (This site's 100-line eval harness tutorial pairs naturally with these techniques: https://everydaydatascience.com/article/agent-eval-harness-catches-regressions)
In favour
The direction of the effect is the same across three independent teams, three different agent stacks, and both enterprise and coding domains. The mechanisms are concrete and checkable: stale state misleads, redundancy wastes, distraction degrades. Nothing here requires believing in a new architecture. It is plumbing.
Against
The honest objections, in order of seriousness:
- Full history is a strawman baseline. No production engineer keeps unbounded history; every real deployment already truncates, summarizes, or windows. The interesting comparison is C3 vs C4 (plain pruning vs pruning plus summary), and the paper's headline 20.6-point figure uses the weaker baseline.
- One study is vendor research on the vendor's own surface. Microsoft researchers, Dynamics 365, MCP. The Claude Sonnet 4.5 cross-check helps, but the domain is still expense itemization.
- "Same performance" in AgentDiet spans a wide band. Input-token savings of 39.9 to 59.7% across LLMs and benchmarks mean the effect is real but setup-sensitive. Your mileage will vary, possibly a lot.
- Peak-token reduction is a narrower claim than cost reduction. ACON's 26 to 54% is about memory pressure. Memory pressure is not the same as the invoice.
What would make me wrong
Three falsifiable predictions, so this piece has an expiry date:
- If next-generation models with genuinely better long-context utilization erase the accuracy gain from pruning, then the effect was a crutch for current models' attention failures, not a law. Watch the pruning-vs-full gap shrink on newer models.
- If the gains fail to transfer to reasoning-heavy work (theorem proving, multi-step math, research synthesis), where each earlier step may be load-bearing, then the claim is domain-bound to tool-use loops with supersedable state.
- If a larger, multi-domain replication finds the savings concentrated below 20%, then the headline numbers were benchmark luck on small task sets. Fifty tasks is not much.
Sources
- Lodha et al., "Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents," arXiv:2606.10209 (June 2026). https://arxiv.org/abs/2606.10209
- Xiao et al., "Reducing Cost of LLM Agents with Trajectory Reduction," arXiv:2509.23586, Proc. ACM Softw. Eng. (FSE 2026). https://arxiv.org/abs/2509.23586
- Kang et al., "ACON: Optimizing Context Compression for Long-horizon LLM Agents," arXiv:2510.00615 (Oct 2025, revised June 2026). https://arxiv.org/abs/2510.00615
Look at your last production agent run. What fraction of its input tokens described a world that no longer existed by the time the agent acted? And what did you pay for them?
Related on Everyday Data Science: Building a 100-Line Agent Eval Harness That Catches Regressions
About the writer
Data Scientist & AI Researcher
Data scientist and AI researcher at Pace University. I coined Artificial Frictional Unemployment, and built the first machine learning model for crop yield prediction in Sierra Leone. Author of Understanding Agentic AI. I write about agentic systems and applied ML, with a bias toward what actually works, and who gets left out when it doesn't.
Found this useful? Passing it on to someone who builds is the best way to help the publication grow.
Built something worth sharing? Write it up for us →