Research questionHow can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?Many VLA evaluations report whether a manipulation task succeeds, but that outcome does not reveal whether failures arise from spatial interpretation, precise execution, or maintaining a multi-step plan. Increasing ambiguity and procedural length make these failure modes harder to distinguish across scenes and embodiments.