Get Started
Research questionHow can we trace which visual, question, or prior-token signals drive VLM generation at each decoding step?Final-answer metrics cannot show how visual input, question text, and already generated tokens shape a response as it unfolds. This obscures whether generation remains visually grounded or increasingly follows the question or its own prior output.
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation FrameworkThe evidence concerns autoregressive VLMs evaluated on image and video understanding tasks, using Qwen3-VL-8B-Instruct and cross-model validation with InternVL2-8B across MAVIS, LLaVA-Video-178K, MiraData, and VLMBias. The supplied results support source-level causal diagnostics without reference answers, including distinguishing prior-driven from visually grounded generations.research paper · Sep 2, 2026
Related questions
How should causal VLMs preserve access to questions placed before image tokens?How can vision-language models correct unsafe generations token by token without disrupting safe reasoning?How can we predict and interpret vision-language model failures to support timely human intervention?How can autoregressive LLM decoding generate multiple tokens in parallel at large batch sizes without sacrificing quality?
Home
Topics
Search
Library