Research questionHow can models segment arbitrary visual concepts precisely from one or a few annotated exemplars without task-specific training?The model must infer which pixels belong to a concept from only a small number of labeled examples. Foreground responses can be coarse or ambiguous, while background content can obscure complete object or part boundaries.