Research questionHow can multimodal models resist harmful intent unfolding across multi-image, multi-turn conversations?A harmful objective can be assembled incrementally across turns and images, leaving each individual exchange seemingly benign. Alignment methods designed for isolated interactions may therefore miss risks that become apparent only from the dialogue history.