Research questionDo newer, larger vision-language models reliably improve autonomous-driving performance without task-specific adaptation?Larger and newer vision-language models often show stronger general reasoning, but that does not necessarily translate into better driving decisions. Autonomous-driving performance can remain inconsistent when models rely on historical actions or struggle to reconcile conflicting visual cues.