Research questionHow can agentic vision-language models acquire and use necessary external evidence without redundant tool calls?Complex image-grounded questions may require visual details or external knowledge unavailable in the initial input. Models may pursue irrelevant evidence or fail to extract useful information from tool outputs, while final-answer supervision does not clearly teach effective evidence acquisition.