Research questionHow can vision-language models express visually grounded answers with context-appropriate information structure?A model may identify the correct visual entity yet present it in a way that conflicts with the surrounding discourse. Content accuracy alone therefore does not establish that a visually grounded answer is linguistically appropriate.