Get Started
Home
Topics
Search
Library
6 min read · Agents · Reinforcement Learning · Added Oct 6 · Paper published Oct 2, 2026

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Source: research paper via Hugging Face Daily Papers
0:00 / 8:24
Agent trajectories collected under one harness poison SFT when you deploy under another—controller prompts and reflection scaffolding become crutches the model won’t have. RSR has the base model rewrite multi-harness successes into clean runbook-guided re-executions under the target harness, lifting pass@3 ~20pp over direct SFT.
TL;DR
Recursive Self-Rewrite (RSR) collects terminal-task successes from several agent Agent harnesses, then has the same base model replan, screen, and re-execute each solution under one general harness, turning harness-specific wins into clean training trajectories that lift pass@3 by up to ~20 pp over direct trajectory SFT.
Why It Matters
You have an agent that runs shell commands to solve tasks (install a package, debug an OpenFOAM simulation, fix a Markdown parser). You notice that the exact same base model solves very different tasks depending on which agent harness you wrap around it. A harness that tracks state carefully nails software tasks; one that reflects after a failed verifier run nails scripting; a vanilla loop nails some exploration-heavy tasks. No single harness covers everything.
The obvious move is to collect all those successful trajectories and fine-tune on them. The problem: each trajectory is stained by its harness. It contains controller prompts, state-tracking commands, and reflection scaffolding that won’t exist at inference time under your general-purpose harness. Train on them directly and the model learns to depend on crutches it won’t have. The paper reports exactly this: direct SFT on raw multi-harness trajectories actually hurts performance on Terminal-Bench 2 (57.0% → 53.4% pass@3) because the model picks up looping and dead-end behaviors from very long reflection trajectories.
How It Works
RSR has three stages, all driven by a single base model, Qwen-3.8-27B.
Stage 1, multi-harness discovery. Run the base model under three harnesses with different control styles: Terminus 2 (plain command loop and the eventual target harness), StateM (explicit workflow states with checked transitions), and RSRT (Terminus 2 extended so that when the verifier fails, the model is told only “failed” and asked to reflect and continue, up to a 180-minute budget). Pool the passing trajectories.
Stage 2, rewrite each success into a clean demo. For each passing source trajectory, the same base model plays three roles:
•
Planner: reads the (compacted) source trajectory and writes a runbook. The runbook captures milestones, checks, recovery tips, and pitfalls, but is prompted not to hand over the finished answer.
•
Critic: sees only the public task description and the candidate runbook, and rejects runbooks that leak the verifier or the solution. Deterministic schema checks run first; the model-as-judge runs second.
•
Executor: in a fresh sandbox under Terminus 2 only, solves the task again using the accepted runbook as private guidance. The runbook never appears in the public trajectory.
They sample K=4 runbooks per source, M=4 executions per runbook, at temperature 0.7. A final filter discards trajectories containing values that couldn’t have been derived from the environment (a sign the runbook smuggled an answer through).
Stage 3, SFT. Keep only the public interaction history from passing rewrites, plus the 766 Terminus 2 passes that didn’t need rewriting, and fine-tune the base model.
for task in tasks: source_trajs = run_under([terminus2, statem, rsrt], task) for traj in passing(source_trajs): for runbook in planner.sample(traj, k=4): if not critic.accepts(task.public, runbook): continue for _ in range(4): new_traj = executor.run(task, runbook, harness=terminus2) if verifier.passes(new_traj) and not leaks(new_traj): sft_data.append(public_only(new_traj))
What They Found
Harnesses are genuinely complementary. Across ~2,929 self-curated terminal tasks, Terminus 2 alone solves 565, RSRT 512, StateM 461. Their union solves 759, a 34.3% relative increase over the strongest single harness. 288 tasks are solved by exactly one harness. This holds up when rollout counts are equalized across harnesses, so it isn’t just a sampling artifact.
Rewriting multiplies usable training data. 2,001 source passes expand to 11,094 verified rewrites covering 975 tasks, with an 86% rewrite pass rate.
RSR beats direct SFT on every benchmark reported. On pass@3, versus Direct SFT: Terminal-Bench 2 53.4% → 74.2%, Terminal-Bench 3 5.4% → 9.5%, Terminal-Bench 4 4.5% → 9.1%, Terminal-Bench Hard 56.0% → 63.0%, Software Terminal 100 3.0% → 6.0%. On Long-Horizon Terminal Bench, process reward goes 0.25 → 0.29, though zero of 46 tasks are fully solved by any variant, so the gain is partial-progress only.
Runbooks genuinely shape behavior. Rewrites sharing a runbook show higher command and command-order overlap than rewrites from different runbooks on the same task (e.g., Markdown: 19.1% vs 12.1% exact overlap). Workflow-level similarity is more task-dependent. A shared runbook raises the odds of similar execution but does not guarantee success; two rewrites from the same runbook can diverge into pass and fail.
Ablation without training: just swapping harnesses on the untrained base model changes Terminal-Bench Hard pass rate from 33.0% (Terminus 2) to 66.0% (RSRT), confirming that harness choice alone moves the needle a lot before any fine-tuning enters the picture.
What’s Useful
•
If you’re training an agent and have several harnesses lying around, don’t pick one, run them all for data collection, then train for the one you’ll actually deploy under. The ~34% coverage lift is the easiest win here.
•
Don’t pool raw multi-harness trajectories into SFT. The paper’s Direct SFT result (regression on Terminal-Bench 2, looping behaviors in the trained model) is a concrete warning. A rewrite-and-re-execute step under your target harness is worth testing before training.
•
The planner/critic/executor split is doable with one model and standard tool-use plumbing. The critic rejecting runbooks that leak the verifier is the piece most teams would skip and most likely to inflate verifier pass rates without real learning, so budget attention there.
•
The released checkpoint is on Hugging Face. Useful if you want to probe how the behaviors transfer, less useful as a drop-in: it’s a 27B base specialized for terminal tasks under Terminus 2.
•
Worth testing: whether the same recipe helps in non-terminal agent settings (browser, code review). The paper only evaluates terminal benchmarks.
Caveats
•
All results use one base model and one target harness. Generalization to other model sizes or harnesses isn’t shown.
•
Several evaluation benchmarks (Terminal-Bench Hard, Software Terminal 100, Long-Horizon Terminal Bench) are self-curated by the authors. Independent benchmarks (TB2, TB3, TB4) show real but much smaller absolute pass rates, especially on TB3/TB4 where even RSR tops out below 10%.
•
Verifier passing is not the same as solution quality: the paper notes verifier success does not imply quality-audit acceptance, and the runbook can still slip hints through despite the critic.
•
The LHTB improvement is a process-reward bump; zero full completions. Don’t read it as “solves long-horizon tasks.”
•
Rollout counts across harnesses were uneven due to sandbox errors (RSRT and StateM had more), so individual-harness rankings in the main table are not a clean comparison. The equal-rollout subset still shows complementarity, which is the load-bearing claim.
Topics
Agents
Reinforcement Learning
Agents
Reinforcement Learning
Up next in Agents
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents223 episodes
Reinforcement Learning109 episodes