Get Started
Home
Topics
Search
Library
Research questionHow reliably do vision-language models correct repeated visually grounded false premises across dialogue turns?A model may initially identify what an image shows but then accept or repeat an incorrect user assumption in later turns. This makes it difficult to distinguish persistent visual grounding from conversational compliance when false premises recur.
AI
Alignment & Safety
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Latest papersRecent research connected to this question, newest first.FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language ModelsThe evidence concerns FPCO-Dialog, which uses 1,080 images and 10,800 question turns in a 10-turn protocol: a correct dialogue prefix is followed by repeated false-premise referring expressions. It evaluates 20 commercial and open-source VLMs with a model-agnostic protocol, the CorrTP@K correction-rate metric, and two independent detectors; results are reported across the benchmark's visual-complexity, object-category, and false-premise strata.research paper · Sep 3, 2026
Related questions
When can frozen vision-language models reliably reason about counterfactual scenes from object-token edits alone?How should causal VLMs preserve access to questions placed before image tokens?How can speech language models consistently use paralinguistic cues in open-ended, multi-turn dialogue?How can we predict and interpret vision-language model failures to support timely human intervention?