Get Started
Research questionHow can multimodal models integrate narrative context with chart evidence to answer multi-step questions?Document questions can require the narrative to identify which entities or values matter before chart evidence can be combined. Models must therefore connect textual constraints with visual data across multiple reasoning steps.
Evaluation & Benchmarks
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense DocumentsThe evidence comes from DocHop, a benchmark of document-style images containing narrative text and charts. Its 2,074 generated examples span six task categories with controllable reasoning depth and visual density; questions use semantic reference labels from the narrative and require multi-chart aggregation. Evaluations cover proprietary and open-source multimodal large language models, with human accuracy above 90% and the best reported model at 62.83%; performance declines as reasoning complexity increases.research paper · Sep 2, 2026
Related questions
How can multimodal models integrate evidence across deeply interleaved text and images?How can multimodal models rely on images or audio rather than language shortcuts?How can multimodal models reason about fine-grained interpersonal relationships from conversational and visual cues?How can long-video QA organize multimodal memory to preserve temporal and cross-modal grounding under limited context?
Home
Topics
Search
Library