Research questionHow can we reliably locate and classify failures in long LLM-agent trajectories?Agent failures can be buried among many interacting steps, making manual inspection costly and LLM-only judging unreliable. Diagnosis must identify both where a trajectory went wrong and what kind of failure occurred.