Get Started
Home
Topics
Search
Library
Research questionWhen does clinical text materially influence pixel-level predictions in medical image segmentation?Multimodal segmenters combine image features with clinical text, but accurate masks may not depend on text consistently across datasets. It is difficult to determine whether text supplies localization evidence or primarily changes global semantic interpretation.
AI
Computer Vision
Health
Image & Video Processing
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence DecouplingThe evidence concerns pretrained vision-language models for medical image segmentation, using text perturbations across BUSI, BTMRI, ISIC, and Kvasir-SEG. Removing text causes major performance drops on BUSI and BTMRI but has marginal effects on ISIC and Kvasir-SEG; the reported influence is mainly global semantic modulation rather than independent spatial localization. An evidential decoder is used to analyze image and text-modulated evidence while maintaining segmentation performance.research paper · Sep 2, 2026
Related questions
How can high-resolution medical image segmentation fuse modalities and clinical text without dense cross-attention costs?How can semi-supervised medical image segmentation prevent appearance variation from corrupting structural cues and pseudo-labels?How can retinal fundus models remain interpretable and accurate when pathology alters vessel appearance?How can medical image models segment lesions despite complex backgrounds and varied morphologies?