Research questionHow can long-video QA organize multimodal memory to preserve temporal and cross-modal grounding under limited context?Long-video QA systems often store captions, frames, transcripts, summaries, and graph facts as separate fragments. At answer time, models must reconstruct which modalities refer to the same event and when it occurred, despite limited context.