Get Started
Home
Topics
Search
Library
Research questionHow can text-promptable video segmentation track targets through disappearance while rejecting visually similar artifacts?Text prompts can identify an object semantically, but video segmentation may lose it when it leaves the field of view, fragment its mask during extreme close-ups, or mistake statues, paintings, and reflections for the target. Such errors can corrupt downstream 3D reconstructions.
AI
Computer Vision
Image & Video Processing
Machine Learning
Multimodal Models
Technology
Latest papersRecent research connected to this question, newest first.ENEAS: Embedding-guided Neural Ensemble for Adaptive SegmentationThe source describes a unified text-promptable system for instance tracking and semantic discovery, using temporal memory for tracking and embedding-based matching with conditional visual-language refinement for ambiguous candidates. Its stated use cases include video, broad libraries, unordered collections, and 3D reconstruction.research paper · Sep 3, 2026
Related questions
How can partially relevant video retrieval locate precise query-relevant moments in untrimmed videos with weak supervision?How can language-guided multi-object tracking preserve identities and context across long-horizon actions in 360° video?How can text-video retrieval preserve temporal structure when matching queries to heterogeneous video frames?How can video continuation models apply new prompts to world states that past frames do not reveal?