Get Started
Home
Topics
Search
Library
Research questionWhich image-text training source best transfers expert ophthalmic knowledge to vision-language models under matched data and evaluation conditions?Ophthalmic image-text training can draw on sources ranging from templated descriptions and clinical reports to broad or highly specialized literature. These sources differ in domain density and dataset composition, making it difficult to determine what drives transfer to clinical tasks.
AI
Computer Vision
Evaluation & Benchmarks
Health
Machine Learning
Multimodal Models
Research Paper
Latest papersRecent research connected to this question, newest first.Scientific Domain Knowledge Improves Vision-Language Fundus ModelsEvidence comes from identical CLIP models fine-tuned on each source, including PubMed-Ophtha: 102,023 image panels with subcaptions from 15,842 open-access PubMed Central articles. Performance is measured across 110 clinical tasks using mean linear-probing AUROC, with controls for fundus-image restriction, image-count matching, and article overlap. The evidence is limited to these models, sources, and evaluations; it suggests domain density as a likely factor but does not establish it conclusively.research paper · Sep 7, 2026
Related questions
How reliably can language and vision-language models answer veterinary clinical questions with retrieval or supervised adaptation?How can vision-language models standardize raw, heterogeneous clinical data before medical inference?How can mammography vision-language models make reliable zero-shot predictions from high-resolution images and homogeneous reports?How can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?