Get Started
Research questionHow can we measure whether long-horizon tool-using agents query hidden state and execute stated plans?In extended tool-mediated tasks, an agent can overlook strategically relevant state that it could query and fail to carry out commitments from its own reflections. Overall outcomes may not reveal these interface-level breakdowns.
AI Agents
Evaluation & Benchmarks
Reasoning
Latest papersRecent research connected to this question, newest first.CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VIThe evidence comes from CivBench, a Civilization VI environment using 76 Model Context Protocol tools and a narration layer that converts visual game state into structured text. Episodes span more than 300 turns and thousands of tool calls; the pilot covers 23 admissible runs across four model families under a shared playbook, characterizing monitoring and commitment execution rather than ranking models.research paper · Sep 2, 2026
Related questions
How can agent decisions be reconstructed for auditing and controlled replay when tool state and authorization context are missing?How can tool-using agents reduce serial action–observation latency without sacrificing task completion?How can text-based world models be evaluated for behavioral fidelity in long-horizon agent planning?How can we train and evaluate LLM agents for tool use across single- and multi-turn workflows with serial or parallel calls?
Home
Topics
Search
Library