Research questionHow can world-model reinforcement learning produce reliable long-horizon driving policies amid interactive traffic and diverse driving styles?World-model reinforcement learning trains policies using imagined rollouts, where small prediction errors can compound over long horizons. Decisions also depend on representing interactions between the ego vehicle and surrounding traffic while accommodating different driving styles.