Get Started
Home
Topics
Search
Library
Research questionHow can mammography vision-language models make reliable zero-shot predictions from high-resolution images and homogeneous reports?Mammography images contain fine-grained visual information that can be lost when they are downscaled for vision-language processing. Reports are also largely dominated by negative or benign findings, making standard image-text contrastive learning less informative for clinical prediction.
AI
Computer Vision
Evaluation & Benchmarks
Health
Multimodal Models
Research Paper
Latest papersRecent research connected to this question, newest first.Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset SamplingThe source concerns mammography vision-language models and reports results for zero-shot density assessment, BI-RADS classification, finding subtyping, and cancer prediction on internal and external benchmarks. It also provides evidence about lesion localization and compares performance with existing open-source mammography and general medical vision-language models.research paper · Sep 2, 2026
Related questions
How can vision-language models reliably report clinically meaningful changes between serial CT scans?How can vision-language models standardize raw, heterogeneous clinical data before medical inference?How can vision-language models maintain visual recognition when modalities are missing and source training data is unavailable?How can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?