Get Started
Home
Topics
Search
Library
Research questionHow can referring expression comprehension localize multiple targets and reject expressions with no valid match in open-world scenes?Standard referring expression comprehension often assumes simple images and exactly one matching object. That assumption makes it difficult to determine whether a model can identify all valid referents or recognize when an expression has no match in complex visual environments.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Machine Learning
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency CheckerThe source introduces OpenRef, a benchmark covering ground and drone views, dark scenes, adverse weather, multi-target and none-target examples, and expressions containing proper nouns, polysemous words, and ordinal terms. It reports F1 for grounding accuracy and Negative Relative Rejection Reliability for rejection against negative expressions.research paper · Sep 2, 2026
Related questions
How can referring multi-object tracking reliably follow every language-matched object across video with low latency?How can reference-guided image generation control distinct attributes of multiple objects in complex scenes?How can multimodal retrieval distinguish correct attribute–object bindings when images share the same concepts?How can video-language models capture the distribution of human interpretations of dynamic facial expressions?