Skills > Agents: Why Powerful LLMs Make Frameworks Obsolete
Yongkang Zou · 2026-04-10 · technical · Agents
Skills > Agents: Why Powerful LLMs Make Multi-Agent Frameworks Obsolete
Hot take: most multi-agent systems built in 2024-2025 will be replaced by a single powerful LLM with good skills and subagent creation.
Not because multi-agent doesn't work. But because frontier models got so good at reasoning, tool use, and self-delegation that the framework layer became overhead. The evidence is mounting.
The Numbers Don't Lie
A recent study compared a single LLM reasoning within one unified context versus multiple LLM instances splitting work and communicating via text, under equal "thinking token" budgets (total tokens used for intermediate reasoning). The result: the single LLM matches or outperforms the multi-LLM setup when computation is normalized.
The gains people attribute to multi-agent architectures? Most are explained by the multi-agent setup simply using more total tokens (more LLM calls, more intermediate reasoning), not by any inherent architectural advantage of splitting work across agents.
The economics are even more damning:
| Metric | Single Agent | Multi-Agent |
|---|---|---|
| Cost per query | $0.005 | $0.016 (3.2x) |
| 10K queries/month | $60 | $192 |
| Debug time (MTTR) | 18 min | 67 min |
| Added latency | baseline | +4.8s per query |
| Accuracy difference | 92.2% | 94.3% (+2.1%) |
A 2.1% accuracy gain for 3.2x cost, 3.7x debug time, and 4.8 seconds of added latency. One enterprise customer was spending $47K/month on multi-agent orchestration that could run on single-agent for $22.7K with 2.1% accuracy loss.
And it gets worse. Unstructured multi-agent networks amplify errors up to 17.2x compared to single-agent baselines. Not 17% worse. Seventeen times worse.
What Changed: Models Got Good Enough
In 2023, you needed CrewAI or AutoGen because GPT-3.5 couldn't handle complex multi-step reasoning alone. You split the work across specialized agents because each agent was dumb enough to need a narrow focus.
In 2026, Claude Opus and Sonnet can:
- Reason through 20+ step plans without losing the thread
- Dynamically create and dispatch subagents when the model decides it needs one
- Use 100+ tools in a single session
- Manage their own context by summarizing, delegating, and pruning
The model itself became the orchestrator. The framework is just while True: think → act → observe.
The Skills Pattern: Let the LLM Compose Its Own Agents
This is what Claude Code's skills system gets right. Instead of pre-defining a rigid agent graph, you give the LLM:
- Skills — reusable instruction sets (markdown files) that teach the model workflows.
/commit,/review-pr,/brainstorm. The model reads the skill and follows it. - Subagent creation — the model can spawn isolated subagents on the fly with custom prompts, tool access, and permission modes. Not pre-configured. Created at runtime based on what the task needs.
- Hooks — event-driven automation (PreToolUse, PostToolUse) that run shell commands. Linting after edits, validation before commits.
- Scheduled tasks —
/scheduleruns prompts on cron. Connect MCP servers (GitHub, Sentry, Slack) and you have autonomous workflows that used to require custom bots or CI pipelines.
graph TB
User[User] --> LLM[Frontier LLM]
LLM --> Skills[Skills]
LLM --> Sub[Subagents]
LLM --> Tools[Tools + MCP]
LLM --> Hooks[Hooks]
LLM --> Schedule[Scheduled Tasks]
Sub --> Sub2[Research]
Sub --> Sub3[Code Review]
Sub --> Sub4[Testing]
Compare this to AutoGen or CrewAI where you pre-define agents, their roles, their conversation patterns, and their handoff logic in code. The Claude Code approach is: tell the model what you want, and it figures out what agents to create.
Token Efficiency: Skills vs Agents
Multi-agent frameworks burn tokens on coordination. Every agent turn requires:
- System prompt (repeated per agent)
- Conversation history (grows with each turn)
- Router/selector deciding who speaks next
- Formatting overhead for structured handoffs
CrewAI consumes nearly 3x the tokens of equivalent single-agent approaches due to its multi-step verification between personas.
The skills pattern avoids this. A skill is a markdown file loaded once into context. A subagent runs in isolation and returns a summary. The parent never sees the subagent's full context. No coordination protocol. No router. Just: delegate, get summary, continue.
Claude Code's /schedule is the ultimate example. Instead of building a custom agent pipeline for "check Sentry every hour and create GitHub issues for new errors," you write:
/schedule "Check Sentry for new errors, create GitHub issues for critical ones" --cron "0 * * * *"
That's the entire "agent system." One prompt, one cron. No framework, no deployment, no infrastructure.
When You Still Need Multi-Agent
To be fair, there are cases where frameworks earn their keep:
- Adversarial validation. Red team / blue team setups where agents genuinely need opposing goals. A single model can't effectively argue with itself.
- Parallel specialized inference. Running 50 classification tasks simultaneously across cheap models. The overhead is computational, not conversational.
- Regulatory/audit requirements. When you need deterministic, auditable agent handoffs for compliance. Skills are too flexible for some regulated industries.
But for 90% of what people build with AutoGen, CrewAI, and LangGraph? A powerful model with skills and subagent creation does the same job with less code, fewer tokens, and faster iteration.
The Prediction
By end of 2026:
- Most "multi-agent" startups will have collapsed their agent graphs into single-model + skills architectures
- The surviving frameworks will be thin orchestration layers (LangGraph-style state machines), not opinionated agent systems
- The term "agent framework" will mean what "web framework" means today: a convenience, not a requirement
/schedule+ MCP + skills will replace 80% of custom agent infrastructure
The best agent system is no agent system. It's a smart model that knows when to delegate.
This site was built entirely with Claude Code using skills and subagent-driven development. No agent framework. Just a powerful model, good instructions, and the ability to create subagents on the fly.
References
- Single-Agent LLMs Outperform Multi-Agent Systems Under Equal Token Budgets — arXiv
- Multi-Agent Orchestration Economics: When Single Agents Win — Iterathon
- The Multi-Agent Trap — Towards Data Science
- Best Multi-Agent Frameworks in 2026 — GuruSup (token cost comparison)
- Claude Code Subagents — Anthropic
- Claude Code Scheduled Tasks — Anthropic
- Claude Code Customization Guide — alexop.dev
- Claude Code Remote Tasks — ComputeLeap