Research questionHow can robotic vision-language-action models generalize across backbones without losing hierarchical manipulation structure?Robotic VLA policies must translate instructions and visual scenes into coordinated manipulation actions. They may overlook relationships among action primitives, become tied to a particular backbone, or optimize objectives at the wrong level of the task hierarchy.