Research questionHow can attention heads be pruned in text-to-image diffusion transformers without losing prompt-specific object identity?During denoising, semantic information may be maintained by structural template tokens and image-to-text interactions rather than by the prompt tokens that initially encode it. This makes it difficult to identify redundant attention computation without disrupting object identity.