Research questionCan multimodal chest-radiograph triage trained on NLP-derived labels reliably match expert severity judgments?Chest-radiograph triage must distinguish urgent examinations from routine ones, but labels extracted from reports may not capture radiologists’ severity judgments. Strong benchmark performance can also coexist with visual explanations that do not localize clinically relevant findings.