Research questionHow can pretrained vision-language-action models reliably perform contact-rich manipulation when goals, scenes, and contacts change?Pretrained vision-language-action models can perform local manipulation skills but often fail when the requested goal or scene differs from their training trajectories. Analytic planners provide compositional control yet remain unreliable for irregular grasps, constrained placements, and articulated-object interactions.