Get Started
Home
Topics
Search
Library
Research questionHow can continual VideoQA learn new tasks without forgetting earlier ones or accumulating task-specific prompts?As a video-language model adapts to successive VideoQA tasks, updates for later tasks can interfere with earlier capabilities. Keeping separate prompts for every task also becomes increasingly costly as the sequence grows.
AI
Computer Vision
Image & Video Processing
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.DynaTokens: Controlling Token Dynamics for Continual Video-Language UnderstandingThe source addresses continual VideoQA with multimodal large language models. It reports evidence from standard continual VideoQA benchmarks, longer domain-incremental sequences, and an ImageQA-to-VideoQA transfer protocol, covering accuracy, forgetting, zero-shot generalization, and cross-modal continual transfer.research paper · Sep 2, 2026
Related questions
How can long-video QA organize multimodal memory to preserve temporal and cross-modal grounding under limited context?How can multimodal models maintain useful visual memory for causal streaming video reasoning under fixed memory?How can pretrained models learn new concepts from evolving streams without task identities or replay while retaining prior knowledge?How can multimodal agents maintain consistent person identities and reason about relationships across long video memories?