Research questionHow can we predict and interpret vision-language model failures to support timely human intervention?Confidence scores and auxiliary classifiers can flag failures without showing which concepts or representations contributed to them. This opacity makes it difficult to assess when a vision-language model needs human review in high-stakes settings.