Get Started
Home
Topics
Search
Library
Research questionHow can multimodal agents maintain consistent person identities and reason about relationships across long video memories?Evidence about one person may be distributed across faces, voices, names, objects, events, and social relations over an extended video history. Systems may recover individual event details while failing to combine those observations into a stable identity profile.
AI Agents
AI Memory
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Latest papersRecent research connected to this question, newest first.ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term MemoryThe evidence comes from ICM-Bench: 839 synthetic clips totaling 141 minutes, with 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark compares caption-memory, memory-augmented, and graph-retrieval systems, with traceable supporting evidence; reported accuracy is 74.0% overall and 60.3% for questions requiring long-term identity profiles.research paper · Sep 8, 2026
Related questions
How can multimodal models maintain useful visual memory for causal streaming video reasoning under fixed memory?How can long-video QA organize multimodal memory to preserve temporal and cross-modal grounding under limited context?How can personalized LLM agents retrieve time-valid memories of persistent and evolving user states?How can multimodal models reason about fine-grained interpersonal relationships from conversational and visual cues?