Get Started
Home
Topics
Search
Library
Research questionHow can agentic vision-language models acquire and use necessary external evidence without redundant tool calls?Complex image-grounded questions may require visual details or external knowledge unavailable in the initial input. Models may pursue irrelevant evidence or fail to extract useful information from tool outputs, while final-answer supervision does not clearly teach effective evidence acquisition.
AI Agents
Computer Vision
Evaluation & Benchmarks
Information Retrieval
Multimodal Models
Reasoning
Latest papersRecent research connected to this question, newest first.Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language ModelsApplies to agentic vision-language models using image cropping, image search, and text search for image-grounded queries. The source reports evidence from seven image-grounded benchmarks within a unified three-tool framework and an 8B-parameter model instantiation.research paper · Sep 3, 2026
Related questions
How can long-video agents choose evidence-acquisition strategies for focused, broad-coverage, or contrastive questions?How can multimodal models rely on images or audio rather than language shortcuts?How can search agents learn when retrieval is necessary and ground answers in evidence without costly supervision?How can image generation and editing agents reliably verify and integrate retrieved multimodal world knowledge?