GraphForge synthesizes training tasks for file-handling agents by crawling real documents, building a cross-file evidence graph, then compiling that graph into both the task prompt and a rubric whose criteria point back to specific files, giving verifiers something concrete to check.
Suppose you’re training an agent to do office work: read a folder of PDFs and spreadsheets, reconcile numbers across them, write a memo. To fine-tune such an agent, you need thousands of example tasks with (a) realistic files and (b) a way to automatically grade whether the agent’s deliverable is correct. These two requirements fight each other.
The paper names the two prior approaches to this problem. EnvCraft generates synthetic files with an LLM and grades with a Python script that inspects workspace state, but the script can’t read inside documents, so a wrong number in a report slides by. NexForge uses real files but ships no task-specific rubrics, so result quality is unchecked. The gap GraphForge targets: real files and reliable per-task verification.
The pipeline has five stages, and the central trick is that the rubric is derived from the files, not written in advance.
1.
Seeds set direction, not content. A seed is an occupation plus a work activity drawn from O*NET, filtered to digitally-executable tasks. Seeds say “a financial analyst does variance analysis” but name no company, no numbers, no dataset. The authors use a greedy coverage rule over (occupation, task type, execution pattern, file family) so the corpus doesn’t collapse onto frequent jobs.
2.
Workspace assembly. A search agent takes a seed, finds a real public case (a company, an incident, a report), and downloads the actual files. Each file gets a hidden role: core, supporting, confuser (plausible but wrong), or ambient (realistic clutter). The agent being trained never sees these labels.
3.
Evidence graph. A model reads the workspace and builds a graph where nodes are facts tied to specific files and edges are cross-file dependencies (“this total in file A should reconcile with the line items in file B”).
4.
Compile graph → task + rubric. The graph is turned into a natural-language task statement plus positive criteria (what the deliverable must contain) and negative criteria (prohibited outcomes). Each criterion carries evidence anchors: pointers to the graph nodes, and hence the files, needed to check it. The working agent sees only the task and the files; the judge sees the anchors.
5.
Rollout, revise, admit. A teacher model (GLM-5.2) attempts the task once. A revision agent then edits the task or rubric if something is ambiguous or unsupported by the files. Only trajectories whose deliverables pass deterministic format checks and score above a rubric threshold are admitted into training data.
for seed in coverage_selected_seeds(onet):
workspace = search_agent.collect_real_files(seed)
graph = build_evidence_graph(workspace)
task, rubric = compile(graph) # criteria carry file anchors
traj = teacher.run(task, workspace)
task, rubric = revise(task, rubric, traj, workspace)
if deterministic_checks(traj) and judge(traj, rubric) > 0.90:
admit(traj)
The scoring formula is a weighted sum: positive criteria contribute weight × fraction_met, negative criteria subtract penalty × violation_degree, normalized by total positive weight. Scores can go negative.
The authors fine-tune two Qwen3.6 base models (27B dense and 35B-A3B mixture-of-experts) on 2,169 admitted trajectories and evaluate on three working-agent benchmarks.
•
GDPval: 220-task gold subset, graded by Elo from pairwise comparisons. The 27B SFT model goes from 1380.0 to 1445.7 (+65.7 Elo) under OpenHands and gains a similar amount under Codex agent harness. The 35B model gains ~101 Elo on both scaffolds.
•
Workspace-Bench-Lite: 27B improves +7.7 points under Claude Code.
•
SpreadsheetBench II: 27B improves +13.7 points; 35B improves +16.5 points.
Gains hold across three agent scaffolds (OpenHands, Codex, Claude Code) even though training rollouts used only Codex, which the authors read as skills transferring rather than scaffold-specific habits.
The more interesting result is the rejection fine-tuning ablation. They take the SFT model, sample 4 rollouts per query on new tasks, and compare three selection rules on 462 matched trajectories:
•
Anchored (judge sees file anchors + verification instructions): consistent gains, +4.0 on Workspace-Bench-Lite.
•
Unanchored (judge sees rubric text but not anchors): smaller or negative gains on the two file-grounded benchmarks.
•
Random-of-4: weakest overall, degrades GDPVal.
On GDPVal specifically, the unanchored arm scored higher in point estimate than anchored, but the authors note no GDPVal RFT differences are statistically resolved at 220 tasks. They explicitly do not claim anchoring wins on GDPVal.
Two audits support the main story. A contamination check finds zero file-hash overlap and no substantive text overlap between training files and GDPVal files; only 13 of 44 GDPVal occupations are covered by the training taxonomy, and SFT gains are actually slightly higher on uncovered occupations. A judge sensitivity probe finds that deleting a cited worksheet drops the target criterion’s score by 0.377 on average while leaving other criteria essentially unchanged (mean |ΔQ| = 0.016), suggesting the judge really does read the anchored evidence. But fine-grained corruptions (wrong numbers, wrong rows) move scores by less than 0.03, which the authors attribute to the judge model’s limits.
•
If you’re building training data for agents that handle messy real documents, the core design pattern is worth borrowing: separate the thing that controls diversity (occupation seeds) from the thing that defines correctness (a graph built after you have the files). It avoids both the “synthetic files feel fake” failure and the “real files but no verifier” failure.
•
The evidence-anchored rubric idea is the specific reusable mechanism. Instead of asking a judge “is this deliverable good,” you give the judge a criterion plus the exact files it needs to check that criterion. The judge-sensitivity ablation gives moderate evidence this helps the judge stay grounded, at least for structural checks.
•
For rejection fine-tuning on self-generated rollouts, the ablation suggests selection rule matters: random-of-k is worse than rubric-guided selection on file-grounded benchmarks. Worth testing in your own RFT setup if you already have a per-task rubric.
•
Data and models are released: HuggingFace collection. Useful if you want to replicate without rebuilding the synthesis pipeline.
•
The pipeline runs on one model, GLM-5.2, everywhere: workspace construction, graph building, task compilation, revision, teacher rollouts, and judging. The judge sensitivity probe shows this model catches structural failures (missing file) but misses fine-grained content errors (wrong numbers in a cited cell). Rubric scores above the 0.90 admission threshold do not mean the deliverable is actually correct in detail, only that structural requirements are met.
•
Scale is small: 2,169 trajectories, and the authors say they have not studied how gains scale with more data.
•
Only two base models, same family (Qwen3.6 27B and 35B-A3B). Cross-family transfer is not demonstrated.
•
The RFT story on GDPVal is unresolved: the unanchored arm looks better in point estimate than the anchored arm, and the authors are explicit that 220 tasks is too few to distinguish them. The claim that evidence anchoring helps RFT selection rests on Workspace-Bench-Lite and SpreadsheetBench II, not GDPVal.
•
Occupation seeds share the O*NET taxonomy with GDPVal, so taxonomy-level overlap exists by construction even though file-level and text-level overlap are zero.