Research questionHow can high-resolution medical image segmentation fuse modalities and clinical text without dense cross-attention costs?Medical images and clinical reports contain complementary spatial, functional, and semantic information. Dense cross-modal interactions become expensive on high-resolution, especially volumetric, feature maps while subtle anatomical details must remain usable.