Get Started
Home
Topics
Search
Library
Research questionHow can scientific-document assistants reason across document modalities while preserving supporting evidence?Scientific questions often depend on relationships among prose, equations, figures, tables, code, and datasets rather than any single modality. Assistants may produce plausible answers without locating or preserving the evidence that supports them, making their usefulness in scientific reading difficult to determine.
AI
Evaluation & Benchmarks
Multimodal Models
Reasoning
Research Paper
Latest papersRecent research connected to this question, newest first.SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document UnderstandingThe source evaluates scientific-document assistants with 124 expert-authored questions spanning seven capability groups, 19 subtasks, and five scientific domains. The 496 instances vary English versus Chinese questions and all-images-first versus interleaved document representations; reported weaknesses include document perception, evidence grounding, verification, and cross-document reasoning. The source also describes a typed evidence-graph representation and associated supervised and reinforcement-learning data, but does not establish performance beyond the evaluated systems and settings.research paper · Sep 4, 2026
Related questions
How can multimodal models integrate evidence across deeply interleaved text and images?How can multimodal medical diagnosis identify informative evidence within each modality without sacrificing accuracy?How can multimodal models integrate narrative context with chart evidence to answer multi-step questions?How can long-video agents choose evidence-acquisition strategies for focused, broad-coverage, or contrastive questions?