Research questionHow can vision-language models interpret whole-body PET/CT by grounding metabolic and anatomical findings?Whole-body PET/CT interpretation requires combining metabolic signals from PET with anatomical and morphological evidence from CT. Existing 3D medical vision-language work has largely focused on regional CT, leaving this multimodal whole-body reasoning problem insufficiently addressed. Latest papersRecent research connected to this question, newest first.MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CTThe evidence concerns MetaStructAtlas, which contains 490 co-registered 3D PET/CT volumes, organ-level segmentation masks, and grounded radiology reports, along with MetaStructVQA, a benchmark of 100,565 grounded visual question-answer pairs. It evaluates state-of-the-art 3D medical vision-language models on anatomical, morphological, and metabolic reasoning.research paper · Sep 3, 2026