Research questionHow can referring multi-object tracking reliably follow every language-matched object across video with low latency?Referring multi-object tracking must connect language expressions to the correct objects and preserve their identities over time. Using multimodal models for this decision can improve language-vision matching but may add latency or produce inconsistent tracks.