Get Started
Home
Topics
Search
Library
Research questionWhen can frozen vision-language models reliably reason about counterfactual scenes from object-token edits alone?Object-token edits change a scene representation without providing a new image, so it is unclear whether a VLM uses the altered objects or relies on learned priors. Evaluation is difficult when post-edit answers have not been annotated.
AI
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It RevealsThe evidence concerns three frozen language-model backbones and remote-sensing scenes from iSAID and VRSBench. An answer-key-free protocol scores logically determined edits and audits them by reversing scoreable choices; findings address explicit edit teaching, token cleanliness and density, and preservation of free-text VQA through the image-free token route.research paper · Sep 3, 2026
Related questions
How can vision-language models express visually grounded answers with context-appropriate information structure?How should causal VLMs preserve access to questions placed before image tokens?How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How reliably do vision-language models correct repeated visually grounded false premises across dialogue turns?