Research questionHow can end-to-end autonomous-driving world models use surround-view inputs without costly future-video generation at inference?World models can predict future scene dynamics for driving, but generating future videos during deployment is computationally expensive. Using only a front camera reduces spatial coverage for maneuvers such as lane changes, merges, and turns.