Get Started
Home
Topics
Search
Library
6 min read · Agents · Evaluation · Added Oct 6 · Paper published Oct 1, 2026

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Source: research paper via Hugging Face Daily Papers
Long-horizon agent outputs (reports, workbooks) have no unit test, so majority-vote and LLM-judge selection fail — 34% of unanimous answers are wrong. VeriHarness splits verification into a disagreement-resolver and consensus-challenger that check claims against workspace files, lifting accuracy ~6 points across five benchmarks.
TL;DR
VeriHarness turns an LLM agent’s own model into a verifier by sampling multiple rollouts, then checking disputed claims and challenging unanimous ones against workspace evidence, lifting single-rollout accuracy by ~6 points across five long-horizon benchmarks.
Why It Matters
Suppose you have an agent that produces a financial report, a multi-sheet workbook, or a code patch. These artifacts contain dozens of individual claims (a revenue figure, a currency, a formula, a required section) that can each be right or wrong independently. In coding or math you can run unit tests or a proof checker, but a report has no equivalent oracle. Today the usual tricks are to sample the model several times and either take the majority vote or ask the model itself to pick the best rollout (LLM-as-a-Judge). The paper’s key empirical finding is that both tricks fail on long-horizon tasks: on APEX-Agents with Claude Opus 4.8, 34% of unanimous answers are wrong, and when rollouts disagree the most frequent answer is correct only 47% of the time, even though 74% of disputed claims contain at least one correct candidate somewhere in the pool. Voting throws that signal away.
How It Works
The idea: don’t just read the rollouts, go check them against the actual files in the workspace, using the same model that produced them. VeriHarness samples N=10 rollouts per task, then splits verification into two parallel investigations run in separate model contexts.
•
The disagreement resolver finds claims where rollouts give different answers, picks the check most likely to distinguish them (framed informally as Expected information gain), runs it against source files, and eliminates candidates the evidence contradicts.
•
The consensus challenger hunts for ways a unanimous answer could still be wrong: wrong units, wrong period in a growth rate, an omitted requirement no rollout addressed. This needs prior knowledge of how artifacts fail, which the harness supplies via a skill library of reusable checking procedures.
A third fresh context then adjudicates both evidence records, picks a base rollout, and emits a revision plan. The verifier never sees the grader, reference answers, or rubrics.
pool = sample_rollouts(task, model, N=10) L_neq = resolver_context(pool, tools, skills) # test disputed claims L_eq = challenger_context(pool, tools, skills) # attack consensus claims base, plan, unresolved = adjudicate(L_neq, L_eq, task, pool) artifact = apply(base, plan) return artifact, evidence_record(L_neq, L_eq, plan, unresolved)
Skills are short text files (sometimes with scripts) like “for a growth-rate label, check that the formula spans the labeled years.” The paper ships 18 human-authored skills covering spreadsheets, PDFs, docs, slides, file bundles, and code patches.
What They Found
Evaluated on five workspace benchmarks (APEX-Agents, Workspace-Bench-Lite, WorkBuddy Bench, SpreadsheetBench 2, JobBench) using Gemini 3.5 Flash and Claude Opus 4.8 as both generator and verifier, with every method receiving the same frozen pool of 10 rollouts:
•
Selection only (pick one existing rollout): VeriHarness beats every baseline on all 10 model-benchmark cells. Average gain over a single rollout is +4.4 points with Flash and +4.1 with Opus. The closest baseline, LLM-as-a-Verifier, lifts the average by about +2.3.
•
With evidence-backed revision (apply the adjudicator’s edit plan): average gain rises to +6.2 points with Flash and +6.4 with Opus, with a standout +11.7 on APEX-Agents under Opus.
•
An agentic verifier given the same workspace and tools but no VeriHarness protocol or skills recovers only about half the gain. So environment access alone is not what matters; the resolver/challenger split and the skills do work.
•
An ablation removing one investigation shows the resolver and challenger recover different errors; running them in one shared context (rather than two) costs ~0.8 points on average.
•
The gain concentrates where it should: on pools where rollouts unanimously agree, selection beats the pool mean by only 0.6 points; on disputed pools, by 6.1 points.
•
Self-evolution: letting the model propose new skills from development-set failures (admitting candidates by dev score, evaluating on held-out tasks) produces a library that beats the human-authored one. Starting from the human library and evolving adds +6.8 on APEX-Agents and +3.7 on SpreadsheetBench 2 over that library.
An Selection oracle that picks the best rollout per task by the grader reaches ~63 average, so a large headroom above VeriHarness remains.
What’s Useful
•
If you run an agent that produces long artifacts and already sample multiple times for a best-of-N judge, the paper’s evidence suggests swapping that judge for a resolver+challenger pattern against your workspace. The gain over LLM-as-a-judge style selection was consistent across five benchmarks and two frontier models, which is unusually broad coverage for this kind of claim.
•
The protocol is training-free and transfers: the authors reproduced most of the selection gain by running the same instructions and skills inside Gemini CLI, Claude Code, and Codex CLI, though JobBench and SpreadsheetBench 2 with Opus lost ground in some CLIs. Worth testing in your own harness before committing to a dedicated runtime.
•
Build a small skill library for your artifact type. The ablation shows removing skills costs the most on APEX-Agents, and the self-evolution results suggest even starting from an empty library and iterating on dev failures gets you useful checks (e.g. “don’t treat a recomputation that matches the pool as confirmation”).
•
Cost matters: the paper reports selection at $1.44/task with Flash and $3.92/task with Opus, cheaper than the pairwise tournament and LLM-as-a-Verifier baselines, because 86-91% of input tokens hit the prompt cache. Revision roughly triples that. These are compute numbers from the paper’s accounting at list price, not deployment benchmarks.
•
The authors release the full pool of ~26,000 rollouts (~$100k of compute) at HuggingFace and the code at GitHub, which lets you reproduce or re-rank without running a model.
Caveats
•
All numbers use N=10 rollouts per task, so you’re paying 10x generation cost before verification even starts. The paper frames this as the standard test-time scaling budget but it’s a real deployment constraint.
•
Verifier and generator share the same model by design. A stronger external judge might do better; the paper deliberately doesn’t test that because it would conflate harness quality with model capability.
•
VeriHarness delivers an artifact plus an evidence record, not a scalar score per rollout, so it isn’t a drop-in reward model for RL.
•
Verification adds multi-turn investigation latency, aimed at reports and workbooks where quality beats response time, not interactive use.
•
Self-evolution requires a development set with grader feedback. The deployment-time verifier still works with no rubrics, but you need rubrics once to grow the library.
Topics
Agents
Evaluation
Google Research
Agents
Evaluation
Google Research
Up next in Agents
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents223 episodes
Evaluation162 episodes
Google Research7 episodes