Get Started
Research questionHow should causal VLMs preserve access to questions placed before image tokens?In causal VLMs, a question before the image can influence visual processing, yet the answer token may have weak access to that question after reading many image tokens. The resulting image-driven answers can ignore what the model was asked.
AI
Computer Vision
Multimodal Models
Latest papersRecent research connected to this question, newest first.Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language ModelsApplies to prompt ordering for visual question answering with open causal VLMs. Evidence covers five open VLMs and multiple benchmarks, including NaturalBench and Winoground, with analyses of attention and answer behavior; claims are limited to the evaluated models and tasks.research paper · Sep 3, 2026
Related questions
How can we trace which visual, question, or prior-token signals drive VLM generation at each decoding step?When can frozen vision-language models reliably reason about counterfactual scenes from object-token edits alone?How can streaming video-language models cut frame-encoding latency while preserving question-relevant evidence?How can vision-language models correct unsafe generations token by token without disrupting safe reasoning?
Home
Topics
Search
Library