Research questionHow can vision-language models ground semantic driving inputs in physically plausible continuous actions with low latency?Vision-language models reason in discrete semantic representations, while vehicle control requires continuous actions constrained by vehicle dynamics. Bridging these representations without introducing trajectory errors or slow sequential generation is difficult.