Research questionHow can text-promptable video segmentation track targets through disappearance while rejecting visually similar artifacts?Text prompts can identify an object semantically, but video segmentation may lose it when it leaves the field of view, fragment its mask during extreme close-ups, or mistake statues, paintings, and reflections for the target. Such errors can corrupt downstream 3D reconstructions.