Get Started
Home
Topics
Search
Library
Research questionHow can computer-vision systems use commonsense to reason about scenes, object relations, and actions?Visual recognition can identify objects without explaining how they relate or what actions make sense in a scene. Commonsense knowledge is difficult to align with visual evidence when datasets are biased, knowledge is incomplete, and representations do not integrate cleanly.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future DirectionsThe source is a survey of knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers for computer-vision tasks. It discusses dataset bias, knowledge incompleteness, integration challenges, cross-modal reasoning, scalable knowledge injection, and neuro-symbolic architectures rather than presenting one proposed system.research paper · Sep 4, 2026VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMsThe source evaluates multimodal large language models on 1,680 questions across 1,249 videos and eight types of visual knowledge, comparing 28 models with human performance. It also reports a visual-knowledge dataset and baseline model, with gains on VKnowU and several other video benchmarks.research paper · Sep 3, 2026
Related questions
How can visual reasoning systems infer and verify formal relational rules from only a few labeled examples?How can robot perception encode action-relevant scene dynamics to improve manipulation generalization?How can robot-learning systems integrate perception, action, and reasoning for reliable long-horizon operation in unstructured environments?How can we measure whether computer-vision models are understandable to independent human evaluators?