Get Started
Home
Topics
Search
Library
Research questionHow can vision-language models interpret whole-body PET/CT by grounding metabolic and anatomical findings?Whole-body PET/CT interpretation requires combining metabolic signals from PET with anatomical and morphological evidence from CT. Existing 3D medical vision-language work has largely focused on regional CT, leaving this multimodal whole-body reasoning problem insufficiently addressed.
AI
Computer Vision
Evaluation & Benchmarks
Health
Multimodal Models
Reasoning
Research Paper
Latest papersRecent research connected to this question, newest first.MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CTThe evidence concerns MetaStructAtlas, which contains 490 co-registered 3D PET/CT volumes, organ-level segmentation masks, and grounded radiology reports, along with MetaStructVQA, a benchmark of 100,565 grounded visual question-answer pairs. It evaluates state-of-the-art 3D medical vision-language models on anatomical, morphological, and metabolic reasoning.research paper · Sep 3, 2026
Related questions
How can vision-language models reliably report clinically meaningful changes between serial CT scans?How can vision-language models infer 3D geometry and temporal continuity from 2D visual observations?How can mammography vision-language models make reliable zero-shot predictions from high-resolution images and homogeneous reports?How can vision-language models standardize raw, heterogeneous clinical data before medical inference?