CrossFit fixes a self-training failure in search agents where a question-writer and answer-solver drift into agreeing on the same wrong answers, by scoring each document’s questions with a separate solver that was never trained on that document.
A Self-evolving agent trains itself without human labels: one copy of the model (the proposer) reads a document and writes a question plus a draft answer, another copy (the solver) tries to answer it, and the proposer gets rewarded based on how often the solver matches its draft. Repeat for a few rounds and you have an automatic curriculum.
The problem the paper names is co-cheating. If the proposer writes a wrong pseudo-label for some document, the solver trains on that wrong label, so later questions from the same document get the same wrong answer, so the proposer’s reward goes up. Internal metrics look great. External correctness stalls or drops. You’ve shipped a model that is confidently wrong in a self-reinforcing way.
The direct baseline is Dr. Zero, a recent proposer-solver loop. The authors instrument it with an independent post-hoc judge (an external LLM auditor that checks answers against source evidence but never feeds back into training) and show that false-agreement mass, the fraction of proposer-solver matches that are matched on the wrong answer, climbs from near zero in round 1 to 6.1% at 4B and 8.8% at 9B by round 3, while apparent reward keeps rising.
The authors try two fixes.
First, the obvious one: verify each question before training on it. Their Multi-sample verification (MSV) samples the same model three times with the source document and three times without it. If both sets of three produce a stable majority answer and the two majorities agree, admit the question and use that majority as the label. This helps a little but not much, because six samples from the same model share the same blind spots. It also costs six extra generations per candidate.
The main contribution is CrossFit, which attacks a different link in the chain: not label quality, but who grades the proposer. Split the source documents into two folds, A and B. Train one auxiliary solver only on questions from A, another only on questions from B. When the proposer writes a question from a document in fold A, score it with the solver trained on B, and vice versa. The grader has literally never seen pseudo-labels derived from that source, so a wrong answer from that source cannot come back as reward through the familiar “solver repeats the error it was trained on” path. The main solver you actually ship still trains on everything; only the feedback signal is cross-fitted.
for doc in sources:
doc.fold = assign_once(doc) # 0 or 1, permanent
for round in range(3):
for doc in sources:
q, draft_label = proposer.generate(doc)
grader = aux_solver[1 - doc.fold] # trained on OTHER fold only
responses = [grader.answer(q) for _ in range(5)]
matches = sum(r == draft_label for r in responses)
reward = (5 - matches) / 4 if 0 < matches < 5 else 0
proposer.update(q, reward)
main_solver.train_on(q, draft_label) # sees both folds
aux_solver[doc.fold].train_on(q, draft_label) # same-fold only
The name borrows the exclusion idea from Cross-fitting in statistics, though the authors note they are not claiming its formal guarantees.
Measured on Qwen3.5-4B and Qwen3.5-9B with the same three-round training schedule:
•
False-agreement mass after round 3: Dr. Zero hits 6.1% (4B) and 8.8% (9B). MSV barely dents it (5.7% / 7.2%). CrossFit cuts it to 3.0% / 3.7%. Combining both reaches 2.0% / 1.7%.
•
Downstream search accuracy, averaged across seven QA benchmarks (Natural Questions, TriviaQA, Long-tail QA benchmarks (PopQA-style), HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle) using Cover-EM: CrossFit improves the average by 8.8 points at 4B and 8.4 points at 9B over Dr. Zero, and by 8.7 / 7.8 points over Search-R1 QA setting. Gains are larger on multi-hop benchmarks (~10 points) than single-hop (~5-7 points).
•
Mechanism isolation: In a controlled replay on 3,000 saved questions where only the grader’s training data changes, source-level splitting drops false agreement to 0.4% / 0.1%. A random split of individual questions (which lets related questions from one document land in both folds) only reaches 5.0% / 6.2%. So the operative variable is really “grader never saw this source,” not just “grader is a different checkpoint.”
•
MSV is largely redundant once CrossFit is in place: adding it on top only buys 0.3 extra points downstream.
The authors are careful: lower false agreement could partly reflect the system rejecting hard questions rather than learning better, so they also report coverage and probe accuracy, which move the right way.
If you are building any loop where a model generates its own training data and another copy of that model grades it, the specific thing to internalize is that the grader’s training-data ancestry matters as much as the grader’s identity. Spinning up a “separate” evaluator does nothing if it was trained on the same pseudo-labels.
Concrete decisions this evidence supports:
•
When you partition data to break feedback loops, split at the source level (document, user, repo), not at the item level. Items from the same source carry correlated errors and will leak across folds if you split naively.
•
Before trusting an internal reward curve from a self-training run, run a post-hoc audit on a sample with an external judge. Track agreement broken down into correct-agreement vs. wrong-agreement. A rising reward with rising false-agreement is the signature of co-cheating.
•
Pure verification-style fixes (sample the same model more times, check consistency) are worth testing but will leave substantial residual error when the samples share a model. Budget accordingly: MSV here roughly doubled the training-run cost for small gains.
•
The CrossFit setup roughly 1.7x’s the training budget because of the two auxiliary solvers. Halving the auxiliary budget preserves most of the benefit, which is worth checking in your own setting.
The paper evaluated open-domain QA with search tools. Extending the same recipe to code agents or tool-use agents is plausible but not demonstrated here.
•
Source exclusion only blocks the direct reuse path. Two solvers can still share errors inherited from pretraining or from documents that overlap in content, and the paper flags “connected sources” as an open problem.
•
The audit uses a commercial LLM judge (the paper names gpt-6-astra/high); the false-agreement numbers inherit whatever biases that judge has.
•
Results cover two model sizes in one family and three training rounds. The authors do not claim lower end-to-end cost, and reserved GPU-hours for the combined method are 2.6-2.7x the Dr. Zero baseline.
•
A drop in false agreement could partly come from the loop admitting fewer hard questions. The paper tries to control for this with coverage and fixed-bank replays, but it is a real confound to watch for when applying the method elsewhere.