Research questionHow can multimodal models maintain useful visual memory for causal streaming video reasoning under fixed memory?Continuous video requires multimodal models to process new visual inputs while preserving information needed for later questions. Strict causality and bounded memory make retaining historical evidence difficult.