Research questionHow can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?Pretrained vision-language-action models can perform broad manipulation but often lack the precision and repeatability required in real-world tasks. Trial-and-error post-training is hindered by unreliable value estimates, policy drift, and the computational cost of running large models.