Get Started
Home
Topics
Search
Library
Research questionHow can we detect and localize failures in long-horizon VLA execution with limited timestamp labels?Long-horizon vision-language-action policies may fail unpredictably, but trajectory-level labels can mark normal pre-failure behavior as failed. Precise localization therefore requires informative timestamp labels without the cost of densely annotating every trajectory.
AI
Alignment & Safety
Machine Learning
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action ModelsApplies to vision-language-action policies for robotic manipulation and detectors that can use action chunks, visual models, or VLA internal representations. The source reports experiments across multiple VLA policies, but does not specify particular robot platforms, datasets, or deployment latency requirements.research paper · Sep 3, 2026
Related questions
How can we reliably locate and classify failures in long LLM-agent trajectories?How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How can vision-language-action systems reliably execute long-horizon manipulation while tracking state and conditional dependencies?How can vision-language-action robots be evaluated for execution quality and decision confidence beyond binary task success?