The Missing Signal: Why Agent Self-Evolution Is Harder Than It Looks

Yongkang Zou · 2026-04-20 · research · Agents

The Missing Signal: Why Agent Self-Evolution Is Harder Than It Looks

There's a wave of AI agent frameworks claiming to "learn" and "self-improve." OpenClaw has 360K GitHub stars. HermesAgent hit 95K in seven weeks. Both promise agents that get smarter over time. But when you look closely at how they learn, a fundamental problem emerges: they're still mostly waiting for you to teach them.

This post unpacks what's actually happening under the hood, why the "worn-in" concept points toward something better, and how the same challenge shows up in an unexpected place: training LLMs for deep research.

The State of Agent Memory

Both OpenClaw and HermesAgent solve the same core problem: LLMs are stateless. Every conversation starts from zero. To make agents useful across sessions, you need persistence.

graph TB
    subgraph OpenClaw Memory
        OC1[MEMORY.md - curated long-term facts]
        OC2[Daily Logs - YYYY-MM-DD.md]
        OC3[SKILL.md files - structured instructions]
        OC4[Memory Compaction - context limit trigger]
    end
    subgraph HermesAgent Memory
        HA1[MEMORY.md - environment facts]
        HA2[USER.md - behavioral model of you]
        HA3[ChromaDB - episodic vector store]
        HA4[Skills - auto-created after complex tasks]
    end

OpenClaw uses a file-based system. MEMORY.md holds curated long-term knowledge loaded at session start. Daily logs capture running context. Skills are directories with SKILL.md files (5,400+ on ClawHub). When approaching context limits, memory compaction forces the agent to write durable knowledge before the window shrinks.

HermesAgent adds a third layer. Beyond MEMORY.md (environment facts) and USER.md (your preferences and communication style), it maintains a ChromaDB vector store indexing every past task execution. What succeeded, what failed, timestamps, semantic similarity search. After complex tasks (5+ tool calls), it autonomously creates skills as structured procedures with pitfalls and verification steps.

The architectural difference matters: OpenClaw's skills are static files you write and maintain. HermesAgent's skills are created autonomously and refined during use. Nous Research reports 40% faster completion on repeated tasks vs. fresh instances.

The Feedback Problem

Here's where both approaches hit the same wall.

graph LR
    subgraph Explicit Feedback
        E1[User corrects agent] --> E2[Agent updates memory]
        E2 --> E3[Better next time]
    end
    subgraph The Problem
        P1[User too tired to correct]
        P2[User doesn't know what to correct]
        P3[Most learning signals are lost]
    end

OpenClaw learns when you explicitly tell it to remember something, or when its agent-evolver meta-skill inspects runtime history for failures. HermesAgent captures implicit signals (accepting output without edits = positive signal, repeated corrections = misalignment). But even "implicit" signals require the user to actively engage with output. In practice, users accept mediocre results because correcting them takes more effort than doing the task differently.

This is the fundamental gap: the agent needs feedback to improve, but the feedback mechanism is a tax on the user.

OpenClaw's self-improving-agent skill and HermesAgent's GEPA (Genetic-Pareto Prompt Evolution) both try to automate this. GEPA uses natural language reflection on execution traces to diagnose why tasks failed, outperforming GRPO reinforcement learning by 6-20% with 35x fewer rollouts. But even GEPA needs a clear success/failure signal to reflect on. When a task half-works and the user shrugs and moves on, the signal is lost.

The "Worn-In" Idea: PokoClaw's Self-Harness

PokoClaw introduces something neither OpenClaw nor HermesAgent really has: a principle called Self-Harness.

The tagline captures it: "Not just remembering what you said. But gradually learning how to work in a way that fits you."

Self-Harness is PokoClaw's answer to the feedback tax. Instead of waiting for explicit corrections or even implicit accept/reject signals, it tracks deeper behavioral signals:

  • Where users feel friction - not what they complain about, but where collaboration slows down
  • Which failures keep repeating - patterns the user stopped bothering to report
  • Which collaboration mistakes should be removed permanently - not appended as preferences, but eliminated as default behavior
  • Which lessons should become future default behavior - promoted from observation to hardcoded patterns
graph LR
    subgraph OpenClaw / HermesAgent
        A[User corrects or accepts] --> B[Agent updates memory]
    end
    subgraph PokoClaw Self-Harness
        C[Agent observes friction] --> D[Identifies recurring patterns]
        D --> E[Transforms into default behavior]
        E --> F[No feedback needed]
    end

This is the "worn-in" concept. Like a pair of jeans that conforms to how you move without you filling out a measurement form, the agent molds itself to your working patterns through sustained use. The key distinction: OpenClaw and HermesAgent's "implicit" signals are still reactive (the agent notices you corrected it). Self-Harness is proactive (the agent notices where collaboration breaks down and fixes it before you complain).

PokoClaw also differs architecturally. It separates work into Main Agent (coordinator), SubAgent (long-lived topic collaboration), and TaskAgent (background execution). This isn't just task management. It's a structure that makes Self-Harness possible. By separating contexts, the system can observe collaboration patterns per-topic rather than in a single chaotic stream.

HermesAgent's USER.md gets partway there. It passively builds a behavioral model across sessions. Meta's PAHF framework formalizes personalization with a three-step loop: seek clarification, ground actions in memory-retrieved preferences, integrate post-action feedback. But both still operate on explicit signals. PokoClaw's Self-Harness aims to learn from what the user doesn't say as much as what they do.

The fully "worn-in" agent would need:

  • Tracking behavioral patterns across hundreds of sessions without explicit tagging
  • Distinguishing between user preference and user laziness (they accepted a bad output because correcting it wasn't worth the effort)
  • Updating without catastrophic forgetting of earlier preferences
  • Converting friction observations into permanent behavior changes, not just memory entries

The Same Problem in Deep Research Training

Here's where this connects to something I've been thinking about from a completely different angle: how we train LLMs to do deep research.

The Benchmark Problem

The standard benchmark for multi-hop reasoning is HotpotQA. 113K questions requiring reasoning across multiple Wikipedia documents. The obvious reward signal is answer accuracy. Did the model get the right answer?

But researchers have discovered this reward is deeply flawed.

graph TB
    subgraph Outcome Reward - The Problem
        O1[Model gets correct answer] --> O2[Reward: +1]
        O3[But did it reason correctly?] --> O4[Unknown]
        O5[Did it use tools well?] --> O6[Unknown]
        O7[Lucky guess vs real reasoning] --> O8[Same reward]
    end
    subgraph Process Reward - The Alternative
        P1[Each reasoning step scored] --> P2[Good search query: +0.3]
        P3[Useful document found: +0.2]
        P4[Correct inference made: +0.4]
        P5[Final answer correct: +0.1]
    end

Why answer accuracy fails as a reward signal:

The evidence is damning. An analysis of GPT-4's multi-hop QA performance found up to 31% of its correct answers are spurious — the model reaches the right answer through flawed reasoning. On comparison questions, the spurious rate hits 46%. CofCA (ICLR 2025) introduced counterfactual data into multi-hop evaluation and revealed a massive performance gap between factual and counterfactual questions, proving models bypass correct reasoning chains entirely. MoreHopQA found that only 38.7% of GPT-4's correct answers achieve "perfect reasoning" where all sub-questions are also correctly answered.

Peng Qi's argument is blunt: stop using HotpotQA. Its extractive format (answers are substrings) is fundamentally mismatched with generative agents. EM/F1 metrics penalize semantically equivalent but differently-worded answers. And the sparse outcome signal collapses all intermediate reasoning quality into a single binary.

Process Rewards: Rewarding the Journey

OpenAI's "Let's Verify Step by Step" (PRM800K) established the foundation: rewarding each correct reasoning step significantly outperforms rewarding only final answers. But the field has exploded since.

graph LR
    subgraph Outcome Reward Model
        O1[Full trajectory] --> O2[Single score]
    end
    subgraph Process Reward Model
        P1[Step 1] --> P1s[Score]
        P2[Step 2] --> P2s[Score]
        P3[Step 3] --> P3s[Score]
        P4[Step N] --> P4s[Score]
    end
    subgraph Generative PRM
        G1[Step] --> G2[Generate verification CoT]
        G2 --> G3[Code-verify if possible]
        G3 --> G4[Score with reasoning]
    end

The generative PRM wave. GenPRM (AAAI 2026) performs explicit chain-of-thought reasoning with code verification before judging each step. A 1.5B GenPRM outperforms GPT-4o. ThinkPRM fine-tunes on just 1% of PRM800K's labels by generating verification chains, outperforming discriminative verifiers across ProcessBench and MATH-500. The insight: a PRM that thinks about why a step is wrong is better than one that pattern-matches.

Credit assignment at scale. GRPO-lambda extends GRPO with eligibility traces for token-level credit, achieving 30-40% training improvement. GiGPO introduces two-level grouping (episode + step) for multi-turn agent tasks, gaining 12%+ on ALFWorld and 9%+ on WebShop without any critic model. BiRM, inspired by A* search, evaluates both backward correctness (was this step right?) and forward probability (can we still reach the answer from here?).

The noise problem. Monte Carlo estimation for process labels is inherently noisy. SCAN (NeurIPS 2025) addresses this with a self-denoising framework where lightweight 1.5B models produce annotations at 6% inference cost of vanilla MC with 39.2 F1 improvement. APRM formulates PRM training as a two-player adversarial game where a Generator creates increasingly deceptive reasoning errors, yielding progressively harder negatives.

What Actually Works for Search Agents

Training agents to search effectively requires different reward signals than training them to reason. Several frameworks have emerged:

graph TB
    subgraph Search Agent Reward Signals
        S1[Answer correctness - sparse, gameable]
        S2[Retrieval relevance - was the doc useful?]
        S3[Information gain - did uncertainty decrease?]
        S4[Search efficiency - over-search vs under-search]
        S5[Tool call quality - well-formed query?]
    end
    S1 -.- S1a[DeepSeek-R1, Search-R1]
    S2 -.- S2a[R3-RAG, RAG-Gym]
    S3 -.- S3a[IG-Search, CoVo]
    S4 -.- S4a[HiPRAG]
    S5 -.- S5a[ToolRM, ToolRL]

Dual-signal approaches. R3-RAG (EMNLP 2025) trains models to interleave reasoning and retrieval via RL with two rewards: answer correctness (outcome) and relevance-based document verification (process). ProRAG (2026) integrates an MCTS-based Process Reward Model into online RL with dual-granularity advantage, resolving credit assignment in multi-hop RAG.

Hierarchical process rewards. HiPRAG introduces hierarchical knowledge-grounded process rewards that address both over-search (wasting tokens on unnecessary retrievals) and under-search (giving up too early). It reduces over-search rate to 2.3% while maintaining high accuracy across 7 QA benchmarks. RAG-Gym provides a systematic framework evaluating prompt engineering, actor tuning, and critic training for agentic RAG, finding DPO most effective for actor tuning with up to 25.6% improvement.

Tool-specific rewards. ToolRL (NeurIPS 2025) is the first comprehensive study on reward design for tool use in RL. Its finding: coarse-grained rewards (answer matching) fail for tool use. Fine-grained rewards that evaluate query formation, output interpretation, and uncertainty reduction significantly outperform SFT. ToolRM builds dedicated reward models achieving 17.94% higher accuracy than frontier LLMs in pairwise tool-use judgments.

The self-play frontier. Absolute Zero (NeurIPS 2025 spotlight) learns to both propose tasks and solve them without any external data, achieving SOTA on math and coding. SPICE generates problems dynamically from corpus interactions via self-play, enabling RL-based reasoning improvement without human curation. CoVo uses consistency and volatility as intrinsic rewards with a curiosity bonus, eliminating the need for external reward models entirely.

An Empirical Reality Check

Not all process signals help. An empirical study on RL for reasoning-search agents found that format rewards improve final performance, but intermediate retrieval rewards have limited impact in some configurations. LLM scale and initialization significantly influence RL outcomes. The choice of search engine critically shapes training dynamics.

This suggests a nuance: process rewards work best when they're meaningful process rewards. Rewarding a model for retrieving a document isn't useful if the model can't distinguish useful documents from noise at its current skill level. The reward has to match the agent's capability frontier, similar to how curriculum learning works.

COMPASS captures this well: it combines Dual-Calibration Answer Reward (outcome) with Decisive Path Reward (process), jointly reinforcing trustworthy consensus answers AND high-quality reasoning chains. The key word is "decisive" — rewarding the steps that actually mattered, not every step equally.

The pattern across all of this: don't reward the destination. Reward the journey. But reward the right parts of the journey.

The Connection

Here's the link I keep coming back to.

Agent self-evolution (OpenClaw, HermesAgent) is stuck on outcome signals. Did the user accept the output? Did the task succeed? These are the agent equivalent of answer accuracy in HotpotQA. Sparse, gameable, and they miss the most important information: how the agent worked toward the result.

graph TB
    subgraph Agent Self-Evolution Today
        A1[Task outcome: success/failure] --> A2[Update memory/skills]
        A3[User feedback: correction/acceptance] --> A2
    end
    subgraph What It Should Be
        B1[Was the reasoning path efficient?] --> B4[Update strategy]
        B2[Were tool calls well-targeted?] --> B4
        B3[Did the agent explore before committing?] --> B4
        B5[Did it ask clarifying questions at the right time?] --> B4
    end

The deep research training literature has already figured out that you need process-level signals. DeepResearcher trains agents end-to-end via RL in real web environments, achieving 28.9-point gains over prompt engineering baselines. Search-R1 uses retrieved token masking with outcome-based reward, improving 41% over RAG baselines. WebRL uses self-evolving curriculum generation from failed attempts.

These training-time techniques point toward what runtime agent self-evolution should look like:

1. Process self-evaluation, not just outcome tracking. An agent should score its own reasoning trajectory. Did I search before guessing? Did I verify my assumptions? Did I use the right tool for this subtask? GEPA's reflection on execution traces is a step toward this, but it still relies on final task success/failure as the anchor.

2. Information gain as intrinsic reward. Borrowing from IG-Search: after each tool call, did the agent's uncertainty decrease? If not, the tool call was wasted. Agents that track their own information gain can learn which queries are productive without any user feedback.

3. Curriculum self-generation. WebRL generates progressively harder tasks based on the agent's current skill level. Agent self-evolution should do the same: identify capability gaps and create practice scenarios, not just wait for users to stumble into edge cases.

4. Worn-in adaptation as implicit process reward. The user's behavior IS the reward signal, but not the way current systems use it. Instead of "user accepted = good output," the signal should be "user spent 2 seconds reading and moved on = low engagement" vs. "user spent 30 seconds reading, copied a section, and asked a follow-up = high value." The depth of engagement is a process-level signal, not an outcome signal.

What's Missing

No current agent framework combines all of these. The closest attempts:

Approach Process Signal Worn-In Self-Curriculum No User Tax
OpenClaw agent-evolver Partial (error logs) No No No
HermesAgent GEPA Yes (reflection) Partial (USER.md) No Partial
PokoClaw Self-Harness Partial (friction tracking) Yes (core design) No Yes
Voyager (Minecraft) Yes (skill verification) N/A Yes (auto curriculum) Yes
WebRL Yes (trajectory scoring) N/A Yes (self-evolving) Yes
ADAS (meta-design) Yes (archive fitness) No Yes Yes

Voyager, WebRL, and ADAS are research systems that work in constrained environments (Minecraft, web arenas). OpenClaw and HermesAgent are production systems that work in the messy real world. The gap between them is the gap between training-time self-improvement (where you can run thousands of rollouts) and runtime self-improvement (where you get one shot per interaction and the user's patience is limited).

Bridging this gap requires solving the same problem from both ends: training agents that evaluate their own reasoning processes (from the deep research side), and building runtime systems that learn from behavioral signals without taxing the user (from the worn-in side).

Where This Goes

I think the next generation of agent self-evolution will look less like "agent remembers what you told it" and more like "agent notices what you never had to tell it." The worn-in concept, properly implemented, is really about building agents with taste: agents that develop an intuition for what you value, not because you explained it, but because they paid attention.

The deep research training community has the right framework. Process rewards, information gain, trajectory-level evaluation, curriculum self-generation. These ideas just haven't crossed over to runtime agent systems yet.

The agent that doesn't need your feedback to improve is the agent that finally deserves to be called self-evolving.

References