Get Started
Home
Topics
Search
Library
Research questionHow can long-video agents choose evidence-acquisition strategies for focused, broad-coverage, or contrastive questions?Long videos contain relevant evidence unevenly: some questions hinge on a localized event, while others require finding occurrences across the recording or comparing explanations. A uniform acquisition strategy can miss necessary evidence before substantive reasoning begins.
AI
AI Agents
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Inference Optimization
Information Retrieval
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video AgentsThe source studies a training-free long-video agent using a shared visual-speech scene index and an acquire–verify–consolidate loop. It reports results on Video-MME-v2, LongVideoBench, EgoSchema, and LVBench with shared query-time models; the supplied evidence indicates gains on several benchmarks and parity on EgoSchema.research paper · Sep 4, 2026VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language ModelsApplies to long-video understanding with frozen vision-language models and executable context constructors. The source examines both weak uniform and stronger hand-crafted initialization and reports transfer to additional long-video benchmarks under matched cumulative visual-token control.research paper · Sep 3, 2026
Related questions
How can long-video QA organize multimodal memory to preserve temporal and cross-modal grounding under limited context?How can research agents refine multi-constraint answers while keeping evidence verified over long horizons?How can long-term conversational QA agents retrieve and reason over temporally dispersed dialogue history?How can long-form RAG answer multi-hop questions across changing temporal, spatial, and relational contexts?