Get Started
Home
Topics
Search
Library
Research questionHow can vision-language models express visually grounded answers with context-appropriate information structure?A model may identify the correct visual entity yet present it in a way that conflicts with the surrounding discourse. Content accuracy alone therefore does not establish that a visually grounded answer is linguistically appropriate.
AI
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.When Discourse Pressures Conflict: Information Structure in Vision-Language Model OutputsThe evidence examines six vision-language models and human participants answering in Hungarian, where Topic and Focus choices are observable through dedicated syntactic positions. It shows how model outputs package information under interacting discourse, grammatical, and definiteness pressures; it does not establish behavior across languages or tasks beyond the supplied setting.research paper · Sep 2, 2026
Related questions
How can vision-language models decide when missing user context requires deferring rather than answering?How should vision-language models answer valid parts of compound queries while withholding unsafe or unanswerable parts?When can frozen vision-language models reliably reason about counterfactual scenes from object-token edits alone?How can multimodal models rely on images or audio rather than language shortcuts?