Get Started
Home
Topics
Search
Library
Research questionHow can language-guided multi-object tracking preserve identities and context across long-horizon actions in 360° video?Limited fields of view can cause targets to leave the frame, breaking identity associations and removing contextual cues needed to interpret long-horizon descriptions. Omnidirectional video retains scene coverage but introduces the challenge of tracking multiple referenced objects across the full view.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Multimodal Models
Research Paper
Technology
Latest papersRecent research connected to this question, newest first.ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object TrackingThe task concerns omnidirectional referring multi-object tracking using 360° imagery and language-guided detection and cross-frame association. Evidence is reported on ORSet, containing 27 scenes, 848 descriptions, and 3,401 annotated objects; the source describes a zero-shot LVLM-driven framework and does not establish broader deployment performance.research paper · Sep 3, 2026
Related questions
How can referring multi-object tracking reliably follow every language-matched object across video with low latency?How can multimodal agents maintain consistent person identities and reason about relationships across long video memories?How can detector-free end-to-end multi-object trackers improve detection without disrupting track association?How can text-promptable video segmentation track targets through disappearance while rejecting visually similar artifacts?