Research questionHow can search agents learn when retrieval is necessary and ground answers in evidence without costly supervision?Sparse outcome rewards do not distinguish useful, evidence-grounded retrieval from redundant searching. Richer process supervision or LLM judging can provide that distinction but adds annotation or inference costs.