AI Agents Drown in Their Own Context: A New Framework Cuts Token Costs

Agentic Context Management: Memory and Cost as Architecture Problems

AI Agents Drown in Their Own Context: A New Framework Cuts Token Costs

Production AI agents often fail not because they reason poorly, but because they can't manage what's in their context: growing histories, large prompts, and bloated tool outputs. This paper argues that treating this as a storage problem is too narrow. Instead, it introduces Agentic Context Management (ACM), a discipline with five primitives—architecting, ingesting, scoping, anticipating, and compacting & consolidation. The authors show that naive context accumulation leads to quadratic token costs, while validated compaction achieves linear costs with preserved fidelity. They present Maximem Synap, a reference implementation that scores 92% on LongMemEval and 93.2% on LoCoMo, and discuss benchmarks' blind spots like latency and context-rot.

Naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity.
  1. nullbio

    Context pollution and rot are probably more important than memory, because facts can usually be retrieved if the agent is good at following breadcrumbs.

    What's also the biggest killer is code rot. Agents are particularly good at death by thousand cuts. They implement something poorly, or incorrectly, or introduce a bad pattern into the project. Then they continue to amplify that badness over time, as they continue to copy from it on subsequent work. It spreads like a virus.

    Keeping these seeds out of the project is very difficult, and cleaning up the rot is very difficult. It also seems like a hard problem to solve because following the existing codebase is something that is good when the code is good, but bad when it is bad. So, seemingly, the solution means more thinking and evaluation for every change that is being made.

  2. samyakk

    ACM, that's the term that I'd been looking for - and your paper explains it clearly. At the end, most of LLM problems are context problems. Getting the correct knowledge into its context window without overpopulating it is the actual engineering effort for most agents. And the solution you present seems promising.

    Both compaction with validation and predictive fetching are the way to go.

    I do not want to write an implementation for this myself, and if Synap is that implementation, I'd like to ask you a few questions:

    1. Does it work with context that's not just agent conversations, but rather documents?

    2. Is it better than RAG on large dataset?

    3. What does on-prem options look like?

  3. respectattentio

    I like to start with memory engineering then reach full system then reducing costs. This allows unlocking full potential of agents.

More from this day

2026-08-26