Get Started
Research questionHow can we reliably locate and classify failures in long LLM-agent trajectories?Agent failures can be buried among many interacting steps, making manual inspection costly and LLM-only judging unreliable. Diagnosis must identify both where a trajectory went wrong and what kind of failure occurred.
AI
AI Agents
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral AbstractionsThe evidence concerns AGENTSCOPE's evaluation on the Who&When and AgentErrata datasets, using structured trajectory representations and LLM-guided reasoning to assess fault localization and failure attribution accuracy.research paper · Sep 2, 2026
Related questions
How can web agents detect impending failure from trajectory prefixes when internal logits are unavailable?How can we detect and localize failures in long-horizon VLA execution with limited timestamp labels?How can LLM-agent systems prevent safety compromises from propagating across workflow boundaries?How can we evaluate LLM flight predictions when accuracy misses safety violations and physical inconsistencies?
Home
Topics
Search
Library