Get Started
Home
Topics
Search
Library
6 min read · Agents · Evaluation · Added Oct 10 · Paper published Oct 8, 2026

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Source: research paper via Hugging Face Daily Papers
0:00 / 8:04
Learn2Play Bench exposes whether LLM agents actually learn from experience via 20 text games with hidden, counterintuitive rules pretraining can’t shortcut. Surprise: keeping raw interaction histories beats distilling them into rules, matched-harness pairings beat cross-pairings cheaply, and top humans still outscore every agent at 84.3 vs 81.8.
TL;DR
Learn2Play Bench tests whether LLM agents actually learn from experience by making them play 20 text games with hidden, often counterintuitive rules that pretraining can’t shortcut. Keeping raw interaction histories often beats summarizing them into rules, and top humans still outscore every agent tested.
Why It Matters
Say you deploy an agent to control an unfamiliar robot API, or to use an internal tool your company just built. Pretraining didn’t cover it. The agent has to try things, watch what happens, and remember what worked. Today we have no clean way to measure whether it’s actually learning versus just recalling training data.
Most existing benchmarks for “self-evolving” agents measure improvement across repeated attempts on familiar tasks like coding issues or web navigation. The problem: if an agent gets better on a familiar task, you can’t tell whether it learned something new from the environment or just surfaced knowledge it already had. The authors cite Reflexion as a canonical early method and Jericho-based benchmarks like J-TTL as closer prior work, but note Jericho games likely leak into pretraining.
How It Works
The authors handcrafted 20 text-based games with hidden rules designed to be novel or counterintuitive. In Poisoner, you serve dishes to a king and must infer his taste constraints from accept/reject feedback. In another game, stealing a colleague’s lunch (not working hard) gets you promoted. The agent is given a goal and a list of valid actions, and must discover rules by acting and observing text feedback.
Key design choices:
•
Deterministic feedback and auto-scoring, so trajectories replay reproducibly.
•
Fixed vs re-shuffled variants: 10 games repeat the same instance each episode (test memory), 10 shuffle characters/layouts while keeping the underlying rule fixed (test transfer).
•
Repeated episodes with no weight updates: the agent gets 5 or 10 episodes per game. What carries over between episodes depends on the harness, which is the scaffolding code that manages tools, memory, and the agent loop.
The benchmark uses four metrics per game: peak score (Max), average (Mean), Learning Gain (last-vs-first episodes), and Learning Slope (linear trend across episodes). The last two are the ones that specifically measure learning, not raw ability.
The experiments compare three axes with the others held fixed:
# axis 1: swap backbone, keep harness = OpenCode for model in [opus_5_5, gpt_6_astra, kimi_k3, ...]: run_all_games(harness=OpenCode, model=model) # axis 2: swap self-evolving method, keep backbone for method in [Naive, Memory, Reflexion, EvoTest, ReasoningBank, AWM]: run_all_games(method=method, model=kimi_k3) # also opus_5, gpt_5_mini # axis 3: swap harness, keep backbone for harness in [OpenCode, ClaudeCode]: run_all_games(harness=harness, model=opus_5)
The self-evolving methods differ in what they carry across episodes. Naive keeps nothing. Memory appends raw histories. Reflexion distills each episode into lessons. EvoTest evolves a guiding prompt and a state-extraction program. ReasoningBank stores reusable reasoning patterns. AWM induces multi-step workflows from successful runs.
What They Found
Raw histories often beat processed memory. With Kimi K3, plain Memory led all four metrics among the self-evolving methods, outperforming Reflexion, EvoTest, ReasoningBank, and AWM. The authors’ explanation: methods that compress episodes into rules can turn an incomplete observation into a definitive constraint. In GemForge, an EvoTest+Opus 5 run tried one gem fusion, failed, and then wrote a prompt telling the agent to avoid other untested combinations, ruling out ones that would have succeeded.
The harness matters as much as the model. Holding Opus 5 fixed, Claude Code beat OpenCode on all four metrics (Max 81.8 vs 74.5). Codex agent harness did the same for GPT-5.6-SOL vs OpenCode. The authors attribute this to models being trained on their native harnesses. Separately, the right harness can be cheaper: Claude Code+Opus 5 beat several self-evolving methods at lower estimated cost per episode.
Top humans still win on peak score. Human Top-1 reached Max 84.3 vs 81.8 for the best agent (Claude Code+Opus 5). Behavioral analysis on fixed games: consecutive-episode action similarity was 0.54 for humans vs 0.66 for agents. After a worse episode, humans hit a new personal best next time 33% of the time vs 22% for agents. Humans explored more; agents exploited sooner.
Changing instances slows learning even when rules don’t change. On re-shuffled games, Sonnet 5 and Gemini Flash models dropped learning slope sharply (e.g., Sonnet 5: control 4.90 → re-shuffled 2.23), while Opus 5 and Kimi K3 lost less. Agents with higher transfer slope recognized why an action worked, not just what worked. Weaker ones repeated episode-specific choices like a menu position after the menu had changed.
Agents can know a rule and still fail to use it. In Haunted Inn, an Opus 5 run correctly stated the room taboos but assigned rooms greedily, leaving no safe room for a later guest. The failure was planning, not knowledge.
What’s Useful
•
If you’re adding memory to an agent for an unfamiliar environment, don’t assume “distill the trajectory into rules” wins by default. The paper’s three-backbone comparison suggests raw retained histories are a strong baseline, especially early when evidence is thin. Worth testing on your own workload before committing to a rule-extraction pipeline.
•
When you can, use the harness the model was trained with. Pairing Opus 5 with Claude Code and GPT-5.6-SOL with Codex improved both score and learning at lower cost than cross-pairings. Not a universal rule, but a cheap thing to A/B.
•
If your agent plateaus, check whether it’s stopped exploring. The paper documents agents concluding a strategy is complete after one failed test (GemForge). Explicit exploration budgets or periodic hypothesis-testing prompts are worth experimenting with. The paper doesn’t evaluate specific fixes.
•
For transfer tasks where surface details change but rules don’t, prefer backbones that showed small LS drops here (Opus 5, Kimi K3) as a starting hypothesis, and watch for the “remembers what worked but not why” failure mode.
•
Benchmark artifacts: project website.
Caveats
All results are on 20 handcrafted text games with 3 trials per configuration. Trial standard deviations are often large (e.g., ±20+ points on individual games), so single-row comparisons can be noisy; the aggregate trends are more reliable than any one cell.
The “rule-discovery” labels in the appendix are qualitative annotations from interaction logs, not independent score measurements. An S label means at least one trial showed complete evidence, not that the agent reliably discovers the rule.
The harness comparison confounds model-harness training alignment with harness design quality; the paper can’t fully separate these. And GPT-6 Astra’s cost figure comes from a confirmed bill while others are standard-price estimates, so the cost frontier mixes measurement types.
Human Top-k groups are selected by achieved score, not independent samples, so the human-agent gap describes top performers, not typical humans (whose Mean Max was 43.6).
Topics
Agents
Evaluation
Reasoning
Agents
Evaluation
Reasoning
Up next in Agents
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
SuperNav: An Agentic Navigation System for Any Task in Any Scene
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents252 episodes
Reasoning125 episodes
Evaluation177 episodes