Research questionHow can we understand VLA latent representations as they evolve across layers during action generation?Vision-language-action models translate visual and linguistic information into robot actions, but their internal states are difficult to interpret. It is especially unclear how spatiotemporal and kinematic information changes across layers during action generation.