Get Started
Home
Topics
Search
Library
7 min read · Agents · Reasoning · Added Oct 8 · Paper published Oct 4, 2026

EVISKILL: Grounding Skill Evolution in Replayable Evidence

Source: research paper via Hugging Face Daily Papers
0:00 / 8:26
EviSkill fixes a quiet failure in agent skill-document evolution: LLM-proposed edits sound plausible but often don’t change behavior, and good edits die inside rejected bundles. The fix is replaying the exact trajectory spans each edit cites under the edited skill — recovering 72.9% of batch-rejected edits.
TL;DR
EviSkill improves the external “skill” documents that LLM agents read at inference time by linking each proposed edit to the specific trajectory spans that motivated it, then re-running those spans under the edited skill to check the edit actually changes behavior as intended before committing it.
Why It Matters
An LLM agent doing multi-step work (calling APIs, running experiments, navigating a text world) often loads a markdown “skill” document: a procedural cheat sheet with action orderings, preconditions, and recovery tips. You don’t retrain the model; you edit the document as the agent runs more tasks and you notice failure patterns.
The standard loop looks like this: run the agent, collect trajectories, ask another LLM to read them and propose edits to the skill, apply the edits, measure validation accuracy, keep the new skill if it scores higher. Prior systems in this space (the paper cites Trace2Skill, SkillGrad, EvoSkill, SkillAdaptor, SkillOpt) all follow that pattern. The authors call this Experience-driven skill evolution.
Two failure modes are endemic to it. First, an LLM that reads a failed trajectory and writes a plausible-looking fix can be wrong. The edit sounds reasonable but doesn’t actually change what the agent does, or it overfits to one task’s quirks. Second, validation is coarse: a batch of edits is accepted or rejected as a bundle, so one bad edit sinks four good ones, and the good ones are discarded.
The authors demonstrate both failure modes empirically before proposing their fix. Running Trace2Skill-generated edits on their source tasks, many edits produce no improvement or active regressions. And when they hand globally-rejected revision batches to GPT-5.5 and let it hand-pick the most promising 1-3 edits, those sub-selections often beat the previous validated skill on training tasks. So the information was there; the all-or-nothing gate threw it out.
How It Works
The core idea: never let an edit into the skill based only on “an LLM read the trace and this edit looks right.” Instead, carry the actual trajectory evidence alongside the edit, and re-execute that evidence under the edited skill to see if behavior actually changes.
The unit of bookkeeping is a Replayable Evidence Card. Each card bundles a proposed correction with pointers to specific (task, epoch, start_step, end_step) ranges in past trajectories that motivated it. When an editor LLM later consolidates cards into concrete skill edits, those pointers come along. Each edit knows exactly which trajectory segments it claims to fix.
Verification works by trigger-range replay. For an edit, the system reconstructs the environment state at start_step by replaying the recorded prefix of the original trajectory (deterministic, not counted as new agent decisions). Then from that state, the agent re-executes under skill + candidate_edit. A judge LLM compares the original segment to the replayed one and returns accept, reflect (edit has a fixable localized problem, get one revision attempt), or reject. Only replay-accepted edits are consolidated into a candidate revision.
That candidate revision then faces a global validation gate on a held-out set. Here’s the second move: if the bundle is rejected globally, each edit gets a post-rejection replay against the previous validated skill (not the full working bundle). Edits that still show local replay support are stashed in a Provisional Edit Ledger. They don’t enter the committed skill, but they ride along in the next epoch’s working skill and get another chance. Unresolved evidence cards also carry forward, and the system generates contrastive cards by comparing an edit’s trajectory this epoch versus last epoch to spot persistent deficiencies.
for epoch in range(E): working_skill = validated_skill + provisional_ledger trajectories = run_agent(training_tasks, working_skill) cards = extract_cards(trajectories) + carried_cards + contrastive_cards edits = editor_llm(group_by_similarity(cards)) verified = [e for e in edits if replay_judge(e, working_skill) == "accept"] candidate = working_skill + consolidate(verified) if val_score(candidate) > val_score(validated_skill): validated_skill = candidate; provisional_ledger = [] else: provisional_ledger = post_rejection_replay(verified, validated_skill)
What They Found
Evaluation covers three interactive benchmarks: AppWorld, ScienceWorld, and ALFWorld, with six action-agent backbones (three GPT variants, three Qwen variants). The evaluator, extractor, and editor LLMs are all GPT-5.5.
•
Overall accuracy. EviSkill wins 14 of 18 model-dataset cells and places top-two in 16. Average gain over a no-skill baseline is 17.93 percentage points. On ScienceWorld with Qwen3.5-9B it beats the next-best skill-evolution baseline by 15.64 points.
•
Evolving can hurt. On ScienceWorld with GPT-5.5, the fixed initial LLM-written skill scores 76.78%, while Trace2Skill, EvoSkill, and SkillGrad score 65.40%, 54.50%, and 67.30%. Several evolution methods actively damage a reasonable starting skill. EviSkill reaches 83.41% and matches or beats the fixed-skill baseline in all 18 cells.
•
Ablations. Removing replay verification drops accuracy by 2.49/4.56/2.85 points on ALFWorld/AppWorld/ScienceWorld. Removing cross-epoch propagation drops AppWorld from 89.68% to 85.52%. Both components contribute, and relative importance varies by environment.
•
Replay catches bad edits. 20.7% of candidate edits on ALFWorld fail initial replay; 41.8% on AppWorld; 42.5% on ScienceWorld. These are edits that the extractor LLM thought reasonable from the trajectories. Many are then recovered through the reflect-and-revise step (49 of 54 reflected edits on ScienceWorld).
•
Rescued edits matter. Of 96 edits provisionally retained after a globally rejected revision, 46.9% enter the validated skill one epoch later and 72.9% are eventually incorporated within the four-epoch budget. These edits would have been thrown out by prior methods.
The authors interpret the pattern as evidence that behavioral verification and cross-epoch retention each solve a distinct failure of experience-driven evolution. The comparison is against reimplemented baselines with the same backbones and task splits.
What’s Useful
•
If you’re building an agent with an external procedural document that an LLM proposes edits to, the key takeaway is simple: re-run the specific trajectory segments the edit cites, under the edited skill, before accepting the edit. You need deterministic replay to reconstruct environment state from a prefix. If your environment doesn’t support that (hosted APIs with side effects, non-replayable state), this exact mechanism doesn’t transfer. The three benchmarks here all support deterministic state reconstruction.
•
Don’t gate edits as all-or-nothing bundles. Even if you’re not building the full card infrastructure, keeping per-edit provenance and letting individually-supported edits survive a batch rejection is worth testing. The paper’s preliminary experiment (hand-picking edits from rejected bundles with GPT-5.5) suggests the signal is there even without a sophisticated retention mechanism.
•
The judge, extractor, and editor are all GPT-5.5 with medium reasoning. Results with weaker judge LLMs are not reported; it’s plausible that replay-based verification depends on the judge being strong enough to distinguish “the agent followed the new rule and it worked” from “the agent happened to succeed anyway.” Worth probing before deploying with cheaper models in those roles.
•
AppWorld shows noticeably more variable gains than ALFWorld or ScienceWorld. If your target domain is closer to tool-heavy API orchestration than embodied or scientific interaction, don’t assume the headline numbers transfer.
•
Code is at GitHub.
Caveats
•
The approach requires an environment where you can save trajectories and deterministically reconstruct state at an arbitrary step. Many real deployments (hosted services, stateful external systems) can’t do this.
•
Evaluation runs four evolution epochs; 27.1% of provisionally retained edits were still unresolved at that budget. Longer runs might change retention and incorporation dynamics.
•
Replay LLM calls are not free. The paper reports workload reductions versus full-trajectory replay (around 66-69% fewer executed steps) but does not report total token or dollar cost against baselines.
•
Local replay acceptance is not the same as global validation acceptance, and the paper is careful to say so. An edit that passes replay can still be excluded by the validation gate, and the mechanism relies on both signals.
•
All judge, editor, and extractor roles use GPT-5.5. The paper does not test whether the method survives with a cheaper judge, which is probably the first ablation anyone deploying this would want.
Topics
Agents
Reasoning
Agents
Reasoning
Up next in Agents
From Evidence to Action: How Tool-Using Agents Fail
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents236 episodes
Reasoning117 episodes