Get Started
Home
Topics
Search
Library
Research questionHow can agents find people across cameras from vague witness clues under spatial-temporal and turn constraints?Witness accounts may be partial or ambiguous, while relevant observations are distributed across camera locations and time. An agent must choose questions and searches before its interaction budget runs out.
AI
AI Agents
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Information Retrieval
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.ARGOS: Who, Where, and When in Agentic Multi-Camera Person SearchThe source describes ARGOS, a benchmark and agent framework with 2,691 tasks across 14 scenarios and Who, Where, and When tracks. Agents access a spatio-temporal topology graph, and experiments with four LLM backbones report turn-weighted success and component ablations; the evidence is limited to these benchmark experiments.research paper · Sep 4, 2026
Related questions
How can multimodal agents maintain consistent person identities and reason about relationships across long video memories?How can language-guided multi-object tracking preserve identities and context across long-horizon actions in 360° video?How can workplace agents interpret human activity traces at the temporal resolution each question requires?How can long-video agents choose evidence-acquisition strategies for focused, broad-coverage, or contrastive questions?