Skills > Agents: Why Powerful LLMs Make Frameworks Obsolete

Yongkang Zou · 2026-04-10 · technical · Agents

Skills > Agents: Why Powerful LLMs Make Multi-Agent Frameworks Obsolete

Hot take: most multi-agent systems built in 2024-2025 will be replaced by a single powerful LLM with good skills and subagent creation.

Not because multi-agent doesn't work. But because frontier models got so good at reasoning, tool use, and self-delegation that the framework layer became overhead. The evidence is mounting.

The Numbers Don't Lie

A recent study compared a single LLM reasoning within one unified context versus multiple LLM instances splitting work and communicating via text, under equal "thinking token" budgets (total tokens used for intermediate reasoning). The result: the single LLM matches or outperforms the multi-LLM setup when computation is normalized.

The gains people attribute to multi-agent architectures? Most are explained by the multi-agent setup simply using more total tokens (more LLM calls, more intermediate reasoning), not by any inherent architectural advantage of splitting work across agents.

The economics are even more damning:

MetricSingle AgentMulti-Agent
Cost per query$0.005$0.016 (3.2x)
10K queries/month$60$192
Debug time (MTTR)18 min67 min
Added latencybaseline+4.8s per query
Accuracy difference92.2%94.3% (+2.1%)

A 2.1% accuracy gain for 3.2x cost, 3.7x debug time, and 4.8 seconds of added latency. One enterprise customer was spending $47K/month on multi-agent orchestration that could run on single-agent for $22.7K with 2.1% accuracy loss.

And it gets worse. Unstructured multi-agent networks amplify errors up to 17.2x compared to single-agent baselines. Not 17% worse. Seventeen times worse.

What Changed: Models Got Good Enough

In 2023, you needed CrewAI or AutoGen because GPT-3.5 couldn't handle complex multi-step reasoning alone. You split the work across specialized agents because each agent was dumb enough to need a narrow focus.

In 2026, Claude Opus and Sonnet can:

  • Reason through 20+ step plans without losing the thread
  • Dynamically create and dispatch subagents when the model decides it needs one
  • Use 100+ tools in a single session
  • Manage their own context by summarizing, delegating, and pruning

The model itself became the orchestrator. The framework is just while True: think → act → observe.

The Skills Pattern: Let the LLM Compose Its Own Agents

This is what Claude Code's skills system gets right. Instead of pre-defining a rigid agent graph, you give the LLM:

  1. Skills — reusable instruction sets (markdown files) that teach the model workflows. /commit, /review-pr, /brainstorm. The model reads the skill and follows it.
  2. Subagent creation — the model can spawn isolated subagents on the fly with custom prompts, tool access, and permission modes. Not pre-configured. Created at runtime based on what the task needs.
  3. Hooks — event-driven automation (PreToolUse, PostToolUse) that run shell commands. Linting after edits, validation before commits.
  4. Scheduled tasks — /schedule runs prompts on cron. Connect MCP servers (GitHub, Sentry, Slack) and you have autonomous workflows that used to require custom bots or CI pipelines.
graph TB
    User[User] --> LLM[Frontier LLM]
    LLM --> Skills[Skills]
    LLM --> Sub[Subagents]
    LLM --> Tools[Tools + MCP]
    LLM --> Hooks[Hooks]
    LLM --> Schedule[Scheduled Tasks]
    Sub --> Sub2[Research]
    Sub --> Sub3[Code Review]
    Sub --> Sub4[Testing]

Compare this to AutoGen or CrewAI where you pre-define agents, their roles, their conversation patterns, and their handoff logic in code. The Claude Code approach is: tell the model what you want, and it figures out what agents to create.

Token Efficiency: Skills vs Agents

Multi-agent frameworks burn tokens on coordination. Every agent turn requires:

  • System prompt (repeated per agent)
  • Conversation history (grows with each turn)
  • Router/selector deciding who speaks next
  • Formatting overhead for structured handoffs

CrewAI consumes nearly 3x the tokens of equivalent single-agent approaches due to its multi-step verification between personas.

The skills pattern avoids this. A skill is a markdown file loaded once into context. A subagent runs in isolation and returns a summary. The parent never sees the subagent's full context. No coordination protocol. No router. Just: delegate, get summary, continue.

Claude Code's /schedule is the ultimate example. Instead of building a custom agent pipeline for "check Sentry every hour and create GitHub issues for new errors," you write:

/schedule "Check Sentry for new errors, create GitHub issues for critical ones" --cron "0 * * * *"

That's the entire "agent system." One prompt, one cron. No framework, no deployment, no infrastructure.

When You Still Need Multi-Agent

To be fair, there are cases where frameworks earn their keep:

  • Adversarial validation. Red team / blue team setups where agents genuinely need opposing goals. A single model can't effectively argue with itself.
  • Parallel specialized inference. Running 50 classification tasks simultaneously across cheap models. The overhead is computational, not conversational.
  • Regulatory/audit requirements. When you need deterministic, auditable agent handoffs for compliance. Skills are too flexible for some regulated industries.

But for 90% of what people build with AutoGen, CrewAI, and LangGraph? A powerful model with skills and subagent creation does the same job with less code, fewer tokens, and faster iteration.

The Prediction

By end of 2026:

  • Most "multi-agent" startups will have collapsed their agent graphs into single-model + skills architectures
  • The surviving frameworks will be thin orchestration layers (LangGraph-style state machines), not opinionated agent systems
  • The term "agent framework" will mean what "web framework" means today: a convenience, not a requirement
  • /schedule + MCP + skills will replace 80% of custom agent infrastructure

The best agent system is no agent system. It's a smart model that knows when to delegate.


This site was built entirely with Claude Code using skills and subagent-driven development. No agent framework. Just a powerful model, good instructions, and the ability to create subagents on the fly.

References