AREX-2 teaches an agent to self-improve at test time by training on long trajectories where an agent iterates for hours on verifiable ML and coding tasks, keeping failed attempts in context so the model learns to recover and keep gaining score as its round budget grows.
Hard agent tasks (training a model, writing a competitive-programming solution, researching an open question) are not one-shot. They need a loop: try, measure the result, revise, try again. In most current agent systems, that loop lives in the scaffold around the model. The harness decides when to retry and what to keep. The model itself just produces one attempt at a time.
That shows up in how training data is built. The standard recipe (used in work like STaR (Self-Taught Reasoner) and rejection-sampling fine-tuning) poses a task, generates attempts, keeps the ones a verifier marks correct, and trains on those. The model sees finished correct answers but never sees how a bad answer became a good one. The intermediate failures, the feedback, and the revisions are thrown away. Trajectories where an agent works for hours and climbs from a weak baseline to a strong solution are almost entirely absent from training sets.
The authors argue this is why agents stall after a few rounds even when given a big budget. They want to move the loop inside the model’s own policy.
The authors decompose self-improvement into two quantities. Reflection sets the average score gain per productive round (call it the mean per-round gain). Long-horizon execution sets how many rounds stay productive before the agent plateaus (call it the effective horizon). Total improvement is roughly those two multiplied together. Reflection without horizon fizzles. Horizon without reflection grinds.
To train both, they build long-horizon improvement trajectories in two domains chosen because feedback is cheap and unambiguous: machine-learning engineering (from GitHub repos) and algorithmic programming (from online judges). A teacher model wraps each source into an environment: a task statement plus a scoring function. An environment is only kept if a reference solution scores well and a naive baseline scores poorly, so there is real room to improve.
Then a strong teacher agent runs in each environment for hours and hundreds of tool calls. It knows the budget, so it establishes a baseline, measures, searches docs, runs small experiments, and revises. The authors call the knowledge it picks up about an environment (which API to call, what the data loader wants) operational knowledge, and some of it is pre-packaged as skills, short documents dropped into the agent’s context.
The key data-selection choice: trajectories are kept or discarded as a whole, based on whether the final score clears a threshold and the transcript is well-formed. Individual bad rounds are not filtered out. Regressions, failed runs, and abandoned approaches stay in. Loss is applied only to the steps that make forward progress after a setback (diagnosing, repairing, submitting an improvement). The failed step itself stays in the context but gets no loss. So the model sees the mess it has to recover from and is trained on the recovery.
for env in environments: # (task, scoring_fn)
traj = []
for t in range(T):
y_t = agent.act(context=traj, skills=skills)
f_t = env.run(y_t) # score + logs/errors
traj.append((y_t, f_t))
if final_score(traj) >= threshold and well_formed(traj):
for step in traj:
loss_mask = step.makes_progress # failures stay, no loss
train(step, mask=loss_mask)
The base model is Qwen3.8-27B. Training mixes these new trajectories with the deep-research data from the prior AREX recipe, unchanged. So any deep-research gain over the previous AREX models is attributable to the new trajectory data, not new research-domain data.
In the training domains. On MLE-Bench-Lite, AREX-2 reaches 81.8 Any Medal (percent of competitions where it earns a medal, averaged over three seeds), the top entry in their comparison table. On FrontierCS, it reaches 70.7, best among open-weight systems reported and within a few points of the strongest closed-weight system. At 27B parameters, it compares favorably with models an order of magnitude larger.
Transfer with no new training data. On deep-research benchmarks, where no newly built trajectories are search tasks, AREX-2 improves over both the 4B and 122B models trained with the previous AREX recipe, hitting 84.0 on BrowseComp, 52.6 on text-only Humanity’s Last Exam (HLE), 92.2 on GAIA, and 93.8 on DeepSearchQA. Same deep-research training data as before, better scores. The authors read this as evidence that long-horizon reflection learned in ML and coding transfers to research.
Round scaling with feedback. On Frontier-CS with a 5-hour budget per problem, AREX-2 climbs from 54.4 at one hour to 65.9 at two hours to 70.7 at five hours, still gaining 2.2 points in the final hour. Two DeepSeek baselines flatline after two to three hours. The effective horizon of AREX-2 is longer.
Round scaling without feedback. On BrowseComp, the agent never sees whether its answer is correct during the run. It uses an inner loop to gather evidence and an outer loop that decides, from its own confidence, whether to accept, refine, or restart. Accuracy rises from 64.8 at 47 turns to 84.0 at 143 turns. The prior AREX (122B) also improves but gains per turn are about a third as large, and AREX-2 passes AREX’s final accuracy with less than half the turns. So the agent is improving by self-judgment, not by being told it was wrong.
Ablation on MLE-bench Lite. Base Qwen3.8-27B scores 28.8. Adding general skills: 41.2. Adding task-specific skills and more rounds: 68.2. Swapping in the trained AREX-2 at the same round budget: 75.8. Trained model with more rounds: 81.8. Operational knowledge, training, and budget each contribute separately. Training alone (same skills, same rounds, M2 vs M4) adds 13.6 points.
•
If you build an agent that iterates against a verifier and it plateaus after a few rounds, the AREX-2 result suggests the fix is in the training data, not the scaffold. Specifically: stop filtering out failed attempts. Keep whole trajectories, apply loss only to the recovery steps. Worth testing on your own domain if you can generate long trajectories with graded feedback.
•
The transfer finding is the surprising part. Training on ML-engineering and competitive-programming trajectories moved deep-research scores without any new research data. If you have one domain with cheap verifiable feedback and another where feedback is expensive, this is evidence (not proof) that long iterative training in the cheap domain can help the expensive one.
•
The ablation separates skills from training. Dropping short domain docs into context (general skills, then task-specific skills) more than doubled the base model’s MLE-bench Lite score before any fine-tuning. If you can’t fine-tune, skills alone are a strong lever.
•
A graded score beats a binary pass/fail for generating useful trajectories, because binary feedback gives the agent nothing to react to on a failed round. If you can design partial-credit scoring for your task, do.
•
Code and models are released: GitHub and HuggingFace.
•
The comparison tables include many model names (GPT-5.6 Sol, Kimi-K3, Claude Fable 5, DeepSeek-V4-Pro, etc.) that do not correspond to publicly known releases at time of reading. Treat the specific headline deltas as comparisons within the authors’ own evaluation setup, not against a reader’s familiar baselines.
•
The MLE-bench Lite number uses skills in context. The ablation shows training contributes 13.6 points on top of skills, but the headline 81.8 is not a training-only number.
•
Trajectory generation needs environments with graded, automatic scoring and a strong teacher agent that can run for hours per task. That is a nontrivial infrastructure cost, and the paper does not quantify it.
•
Transfer is demonstrated from (ML engineering, coding) to (deep research). The claim that long-horizon reflection is fully domain-agnostic is a hypothesis the results are consistent with, not something the paper proves across arbitrary domains.
•
A gap remains against the strongest reported BrowseComp and HLE systems, so this is not a universal frontier result, it is a strong result at 27B parameters.