Context Rot Is Real, and Subagents Are the Fix
Yongkang Zou · 2026-02-10 · technical · Agents
Context Rot Is Real, and Subagents Are the Fix
Working on multi-agent systems at Epiminds, I spent a lot of time thinking about why agents get worse the longer they run. The answer is context rot, and the most practical solution I've found is subagent-driven context isolation.
This post breaks down the problem, the architecture, and what I learned building production agent systems.
The Problem: Your Agent Gets Dumber Over Time
Every LLM suffers from context rot. Chroma's research tested 18 frontier models (GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, Qwen3-235B) and found that every single one gets measurably worse as input length increases.
Three mechanisms are at play:
- Lost-in-the-middle effect — models attend well to the start and end of their context but poorly to the middle. Accuracy drops 30%+ when relevant info sits in positions 5-15 of a 20-document window.
- Attention dilution — transformer attention is quadratic. At 100K tokens, the model tracks 10 billion pairwise relationships. Each new token dilutes attention on everything else.
- Distractor interference — semantically similar but irrelevant content actively misleads the model. This is especially bad for coding agents that accumulate search results and backtracking noise.
The practical effect: Cognition found that by the time a coding agent finds the relevant code, the signal-to-noise ratio in its context is 2.5%. Doubling task duration quadruples the failure rate.
graph LR
A[Fresh Context] --> B[Noise Accumulates]
B --> C[Context Rot]
C --> D[Agent Failure]
The Solution: Subagent Context Isolation
The key insight: don't let one agent accumulate everything. Delegate work to subagents that run in isolated context windows, do their job, and return only a summary to the parent.
This is the same principle behind Claude Code's subagent architecture, LangChain's Deep Agents, and what we built at Epiminds. The parent orchestrator stays clean. The heavy lifting happens elsewhere.
graph TB
O[Orchestrator] -->|task| S1[Data Analyst]
O -->|task| S2[Copywriter]
O -->|task| S3[Research Agent]
S1 -->|summary| O
S2 -->|summary| O
S3 -->|summary| O
Without isolation, the orchestrator would hold 180K+ tokens of accumulated context. With isolation, it holds maybe 2K tokens of clean summaries. The quality difference is dramatic.
How We Built This at Epiminds
At Epiminds, we built a multi-agent marketing platform with 20+ specialized agents coordinating under a supervisor architecture. The core product, Lucy, manages campaigns end-to-end: reporting, pacing, creative analysis, budget optimization.
The framework follows a supervisor pattern similar to AutoGen's GroupChat and VoltAgent's agent orchestration, but tuned for marketing workflows.
The Supervisor Pattern
sequenceDiagram
participant User
participant Lucy as Lucy (Supervisor)
participant DA as Data Analyst
participant SA as Strategy Agent
participant CW as Copywriter
participant CA as Creative Analyst
User->>Lucy: Optimize Q2 campaign
Lucy->>DA: Analyze performance
DA-->>Lucy: CTR down 12%, mobile 2x desktop
Lucy->>SA: Recommend reallocation
SA-->>Lucy: Shift 30% to mobile
Lucy->>CW: Generate mobile copy
CW-->>Lucy: 3 variants ready
Lucy->>CA: Score variants
CA-->>Lucy: Variant B wins
Lucy-->>User: Action plan + creatives
Each agent gets only what it needs. The Data Analyst never sees brand guidelines. The Copywriter never sees raw API metrics. This isn't just about context window limits. It's about signal quality. A copywriter with 60K tokens of campaign metrics in its context writes worse copy, even if the window can hold it.
Why Not Just Use a Bigger Context Window?
Because the problem isn't capacity. It's attention.
The AgentRM paper identifies two failure modes in agent systems: scheduling failures (system unresponsiveness) and context degradation (agent "amnesia" from unbounded memory growth). Both get worse with larger windows, not better.
| Approach | Context Size | Signal Quality | Failure Rate |
|---|---|---|---|
| Single agent, 200K window | 200K tokens | Degrades over time | High after 35min |
| Multi-agent, shared context | 200K shared | Polluted by all agents | Medium-high |
| Subagent isolation | 2-5K per parent turn | Stays clean | Low |
The Pattern in Practice
Whether it's Epiminds' marketing agents, Claude Code's subagent framework, or AutoGen's v0.4 actor model, the pattern is the same:
flowchart LR
subgraph BAD[Shared Context]
A1[Agent A] --> SC[Growing\nShared Pool]
A2[Agent B] --> SC
A3[Agent C] --> SC
SC --> R1[Noise]
end
subgraph GOOD[Isolated Subagents]
O2[Orchestrator] -->|task| B1[Agent A]
O2 -->|task| B2[Agent B]
O2 -->|task| B3[Agent C]
B1 -->|summary| O2
B2 -->|summary| O2
B3 -->|summary| O2
end
The rules are simple:
- Orchestrator holds summaries, not raw data. It decides what to do, not how to do it.
- Each subagent gets a focused brief. Include what it needs, exclude what it doesn't.
- Subagents return structured summaries. Not their full chain-of-thought. Not intermediate results. A clean answer.
- Parallel when possible. Independent tasks run concurrently. The orchestrator waits for all, then synthesizes.
What This Means for Your Agent Architecture
If you're building multi-agent systems and your agents degrade after 10-15 turns of conversation, the fix isn't a bigger context window. It's isolation.
The Intrinsic Memory Agents paper calls this "structured contextual memory" — each agent maintains memories specific to its role, ensuring heterogeneity rather than a shared soup of context.
Context engineering is becoming as important as prompt engineering. The question isn't "what do I put in the prompt?" It's "what do I keep out?"
Built with this approach at Epiminds (marketing agent platform, $6.6M seed from Lightspeed). Currently applying similar patterns to yongkang.dev — where Claude Code's subagent-driven development helped ship this entire site.
References
- Context Rot: Why LLMs Degrade as Context Grows — Morph/Chroma research
- AutoGen: Multi-Agent Conversation Framework — Microsoft Research
- LangChain Deep Agents — MarkTechPost
- Claude Code Subagents — Anthropic
- AgentRM: OS-Inspired Resource Manager for LLM Agents
- Intrinsic Memory Agents
- Context Discipline and LLM Performance
- AutoGen v0.4 Introduction — Victor Dibia
- Epiminds + VoltAgent Case Study