Research questionHow can vision-language models decide when missing user context requires deferring rather than answering?A response that is generally reasonable can be unsafe for a particular user when relevant medical, emotional, or situational context is unavailable. Visual information may also enter the model’s text representation early and suppress textual signals that should prompt caution.