OneSearch-VL trains a single vision-language agent to do deep web research across single images, image sets, and videos by anchoring every claim to a Visually Grounded Evidence Graph that ties visual regions to source-backed facts, with a rubric reward that scores whether the agent actually used the right visual evidence.
Suppose you want an agent that can answer “which car in frame 12 is made by a company Volkswagen owns?” The agent has to find the right frame, crop the right object, reverse-image-search it to a real-world entity, then look up ownership facts on the web. Today’s multimodal search agents usually specialize: some handle single-image queries (MMSearch-style), others handle video deep research. Training one policy for all three input types (single image, multi-image, video) is awkward because the supervision signals differ, and because final-answer rewards don’t tell you why a trajectory failed. Did the agent miss the object? Pick the wrong frame? Retrieve the wrong Wikipedia page? Hallucinate a fact the retrieved page didn’t support?
The closest prior baseline here is OpenSearch-VL, which trains a single-image search agent using answer-correctness and query-quality rewards. OneSearch-VL extends that recipe to multi-image and video, and adds process-level supervision tied to evidence structure.
The core data structure is a Visually Grounded Evidence Graph (Visually Grounded Evidence Graph): for each training question, it records which visual regions (anchors) point to which real-world entities, which source URLs and quoted snippets back each fact, and which composition operation (compare, count, arithmetic, multi-hop lookup) produces the final answer. This graph is built once by a data pipeline and then reused as the ground truth for task construction, expert trajectory filtering, and reward computation.
The data engine starts from 2.5M YouTube videos, filters to 70k by metadata and LLM screening, then for each video builds a hierarchy of events, key frames, and localized objects. Each object is routed through reverse image search or OCR-plus-text-search to identify real-world entities and attach source-supported facts. From this input-level evidence graph, the engine generates compositional questions, verifies them against the evidence, and rewrites explicit entity names into visual references (“the blue car in image 2” instead of “the Ford Mustang”). An expert model (Seed 2.0 Pro) then solves each question in the real tool environment, and trajectories that pass answer-correctness and process-quality judges become OneSearch-VL-SFT-110K (~110k trajectories, roughly balanced across the three input types).
After Supervised Fine-Tuning on Qwen3-VL-8B, Reinforcement Learning uses a composite reward. The novel piece is the Evidence-aware Visual-Grounded Rubric reward (Evidence-aware Visual-Grounded Rubric reward), which runs two separate judge calls per rollout: r_trace checks whether tool observations actually establish the required facts and whether the reasoning follows those observations without unsupported substitutions; r_ground checks whether the correct visible objects, regions, or video frames were identified and used to drive subsequent searches. These combine with answer-correctness and query-quality rewards (weights 0.6 / 0.2 / 0.2 for answer / query / EVGR) and feed into Group Relative Policy Optimization (GRPO) optimization, with the loss applied only to policy-generated tokens, not tool observations.
# Per training question:
vgeg = build_vgeg(visual_input, question) # anchors, facts, ops, deps
for rollout in sample_group(policy, question, G=8):
traj = run_in_tool_env(rollout) # crop, OCR, image/text search
r_acc = judge_answer(traj, reference)
r_query = judge_query_quality(traj)
r_trace = judge_evidence(traj, vgeg) # facts supported by observations?
r_ground = judge_grounding(traj, vgeg) # right regions/frames used?
R = fmt_valid(traj) * (0.6*r_acc + 0.2*r_query + 0.2*(r_trace+r_ground))
grpo_update(policy, rewards)
On two new operation-organized benchmarks the authors built (OneSearch-MI-Bench for multi-image, OneSearch-Video-Bench for video), the 8B model beats the Qwen3-VL-8B agent baseline by 20.2 and 17.6 points respectively, and by 27.0 points on the external VideoDR benchmark. On seven established single-image deep-research benchmarks (SimpleVQA, VDR, MMSearch, LiveVQA, BrowseComp-VL, FVQA, InfoSeek), the average is 58.3 versus 42.0 for Qwen3-VL-8B and 56.6 for OpenSearch-VL-8B. Gains concentrate in compositional operations: on multi-image tasks, knowledge-conditioned counting improves by 29.8 pp and multi-anchor arithmetic by 26.3 pp.
The ablations are where the structural claims get support:
•
Joint training across input types helps, doesn’t hurt. Training on video-only trajectories gives 55.1 average across six benchmarks; adding all three types reaches 55.8. Video-only training also improves single-image benchmarks, suggesting the trajectories transfer.
•
EVGR adds signal beyond answer and query rewards. Starting from the joint SFT model at 55.8, answer-only RL reaches 56.3, adding query quality reaches 57.3, and adding both EVGR dimensions reaches 61.1. Each process dimension alone (trace or ground) lands around 59, and they’re complementary.
Note the authors judge final answers with GPT-4o as a binary correctness decision, so “accuracy” here is LLM-judged, not exact match.
•
If you’re building a multimodal retrieval agent and getting opaque failures, the EVGR split (evidence traceability vs visual grounding) is a diagnostic worth borrowing even without their training setup. Running two targeted judge calls on your rollouts tells you whether failures are coming from wrong-object-identified or wrong-fact-retrieved, which answer-match scores can’t distinguish.
•
The VGEG construction pipeline is heavy (it uses Seed 2.0 Pro, GPT-4.1, Qwen3-VL-30B, and external image+text search services across five stages). If you want to replicate the training, budget for that; if you just want the benchmarks, the operation-organized evaluation design is reusable with lighter infrastructure.
•
The paper does not release code or weights in the supplied text, so treat this as an architecture-and-recipe reference rather than a plug-and-play system. The authors mention the agent depends on external search tools and changing webpages, so even their own numbers aren’t exactly reproducible without retrieval snapshots.
•
Worth testing: using a VGEG-style structured rubric to score agent trajectories in your own domain (not just multimodal search) where “did the agent use the right evidence” is distinct from “did it get the right answer.”
The two headline benchmarks (OneSearch-MI-Bench, OneSearch-Video-Bench) are introduced by the same authors who built the training data pipeline, using the same VGEG structure that supervises training. Gains on these benchmarks partly reflect alignment between training and evaluation design. The external VideoDR gain (27.0 pp over Qwen3-VL-8B) is the strongest third-party evidence that the method generalizes.
All scoring is LLM-as-judge (GPT-4o for final answers, Seed 2.0 Pro for EVGR), inheriting those models’ biases. The qualitative cases in the appendix show the agent still produces inconsistent intermediate facts (e.g., conflicting building heights) even when the final answer is correct, so evidence traceability improves but is not solved. Finally, the comparison is primarily against Qwen3-VL-8B with tool access and OpenSearch-VL-8B; the paper does not claim to beat larger closed models uniformly.