Research questionHow can mammography vision-language models make reliable zero-shot predictions from high-resolution images and homogeneous reports?Mammography images contain fine-grained visual information that can be lost when they are downscaled for vision-language processing. Reports are also largely dominated by negative or benign findings, making standard image-text contrastive learning less informative for clinical prediction.