Research questionHow can multimodal models jointly learn region captioning and spatial localization without text annotations?Region captioning must produce descriptions that distinguish visual content, while localization must recover the corresponding spatial region. Jointly learning both capabilities is difficult when textual annotations are unavailable.