Get Started
Home
Topics
Search
Library
Research questionHow can a single search agent improve multi-hop web research without sub-agents or test-time verification?Multi-hop web research requires an agent to connect evidence across several pages while retaining the information needed for later steps. Long search trajectories can overwhelm the available context, making both training and reliable performance difficult without additional agents or verification passes.
AI
AI Agents
Evaluation & Benchmarks
Inference Optimization
Information Retrieval
LLM Pretraining & Post-training
Machine Learning
Natural Language Processing
Reasoning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Iris: Climbing to the Search FrontierThe source studies Iris-mini and Iris-pro, trained through constructed hyperlink-graph questions, filtered search trajectories, supervised fine-tuning, and reinforcement learning with live search. It evaluates a single ReAct agent with and without context management on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE, using fixed tools, context limits, and a judge; the reported results are for these systems and benchmarks.research paper · Sep 3, 2026
Related questions
How can search agents learn when retrieval is necessary and ground answers in evidence without costly supervision?How can research agents refine multi-constraint answers while keeping evidence verified over long horizons?How should scientific agents be evaluated on underspecified, attachment-rich requests without ground truth?How can search agents retrieve relevant posts from heterogeneous social feeds for complex queries?