Research questionHow can text-video retrieval preserve temporal structure when matching queries to heterogeneous video frames?Videos contain changing appearance, motion, and semantic transitions that a single pooled representation can obscure. This makes it difficult to align a natural-language query with the relevant temporal content.