Predictive Credit asks whether an AI research agent’s written explanation of a planned experiment actually helps forecast that experiment’s outcome, by comparing matched explanations against description-only and donor (someone else’s) explanations under the same forecaster and outcome. Across 336 prospective states, the matched explanation showed no confirmed predictive gain.
Autonomous research agents now propose experiments, run them, and write up why they expected the result. Benchmarks today score the result. Nobody scores the explanation. If an agent writes “stronger regularization will close the train-test gap because X,” that rationale is either genuinely informative about the outcome or decorative text. Treating all explanations as equally valuable rewards fluent storytelling over actual insight.
The authors want a measurement protocol, not a new agent. Prior simulatability work (notably Hase on Leakage-adjusted simulatability) scored explanations by whether they help an observer predict a model’s output. This paper pushes the same idea onto predicting executed experimental outcomes, which is what research agents are supposed to be good at.
The setup is a paired forecast. For each experiment i with public state x, planned intervention a, explanation h, and realized outcome y, a fixed forecaster LLM sees one of three contexts:
•
Description D = (x, a). The baseline: facts and the planned change, no rationale.
•
Matched M = (x, a, h). Adds the agent’s own explanation.
•
Donor S = (x, a, h from another case). Adds somebody else’s explanation, grabbed from a different experiment under a frozen shuffle.
The gain of matched over description measures whether the explanation adds anything at all. The gain of matched over donor tests whether the gain is specific to this case, or whether any plausible-sounding rationale would do. Both use the same outcome and loss, so differences come from the explanation text.
Explanations are elicited as a prediction card with six slots: direction, quantitative target, intermediate observable, benefit regime, falsifier, counterfactual. This structure makes commitments testable. Each forecast returns a point estimate and an 80% interval, scored with Interval score, along with repeat drift between two independent calls.
Five checks are tracked together: commitment (did the card get filled), delivery (did evidence-backed content reach the forecaster), predictive value (D vs M), alignment (S vs M), and sensitivity to known signals.
for state in states:
for context in [D, M, S]:
y_hat_1, interval_1 = forecaster(context) # two independent calls
y_hat_2, interval_2 = forecaster(context)
drift[context] = abs(y_hat_1 - y_hat_2)
err[context] = mae([y_hat_1, y_hat_2], y_true)
iscore[context] = interval_score(interval_1, interval_2, y_true)
gain_desc = R(D) - R(M) # matched vs description
gain_align = R(S) - R(M) # matched vs donor
Studies run on three settings: a 120-state controlled v5 learning environment, Tox21 with 72 states across 12 assay endpoints, and 24 OpenML tabular tasks giving 144 states. The forecaster is DeepSeek-V4-Pro, with a follow-up DeepSeek-V4-flash replay.
The headline is a null, honestly reported. The frozen Tox21 primary test (does the matched card lower ROC AUC Interval score by at least .005) failed: estimate −.0026, 95% interval [−.0174, +.0104]. OpenML’s joint rule on card formation, point-equivalence, and repeatability also failed. v5 was inconclusive. Donor-vs-matched intervals on Tox21 and OpenML spanned zero. So matched explanations did not reliably beat donor explanations on predictive accuracy.
What did move: structured cards sharply cut repeat drift. On Tox21 Pro, drift fell from .01004 (description) to .00357 matched and .00411 donor, reductions of 64.5% and 59.1%. A Flash replay reproduced that drift reduction but raised matched point MAE from .01823 to .02020 and missed interval-score equivalence. Translation: cards make the model self-consistent across calls without making it more accurate.
Elicitation changes commitment rates massively. v5 produced complete six-slot cards in 0/60 spontaneous proposals vs 59/60 when explicitly asked. On OpenML, 78/144 elicited cards came back as empty scaffolds, and assigning the full card widened 80% intervals by 21% while coverage dropped slightly (49.3% vs 51.4% for description).
Direction accuracy mostly came from a sign prior: v5 agents predicted “improvement” in 109/111 cases, scoring 58.6% vs a constant-positive baseline of 57.7%. The forecaster isn’t really discriminating; it’s betting on the prior.
Two positive controls show the forecaster can use strong information when present. A researcher-authored note disclosing the latent functional form in a quadratic-feature experiment cut point MAE by 2.60 pp vs description and 4.36 pp vs a deliberately false mechanism note. A separate Exact-signal calibration where the true outcome was handed in as a hint achieved Spearman .99996. So the pipeline detects real signal; the agent explanations just don’t carry much.
•
If you’re building a research-agent benchmark, this gives you a protocol to score rationales alongside results. The paired D/M/S design with a fixed forecaster and loss is the reusable piece. You do not need the specific Tox21 or OpenML tasks.
•
Separate repeatability from accuracy in your own evaluations. This paper shows a case where structured prompting cuts call-to-call drift by ~60% while leaving point error flat or slightly worse. Reporting only drift would overclaim.
•
Treat agent-written “direction” claims with suspicion. Compare to a constant-direction baseline before celebrating any direction-accuracy number. The sign prior absorbed most of the signal here.
•
Worth testing in your own setup: elicited cards may improve downstream auditability (explicit falsifiers, counterfactuals) even when they don’t improve forecasts. The paper measures prediction, not usefulness for humans reviewing a proposal.
•
The positive controls suggest a diagnostic: before trusting that an agent’s natural explanation is informative, inject a hand-authored mechanism note on the same tasks and confirm the forecaster responds to it. If your pipeline fails that check, no explanation experiment will work.
The null is specific to this forecaster, these tasks, and these donor resolutions (within-cell seed swaps on Tox21 and OpenML, cross-intervention on v5). The paper is careful to call natural-explanation credit “unconfirmed at the tested donor resolutions,” not refuted. A different forecaster or coarser donor pool could change the picture.
Delivery was weak on OpenML: more than half of elicited cards were empty scaffolds after extraction, so the D-vs-M comparison is partly a test of whether cards even arrived with content. The paired-content subset is only 26 states and the authors flag it as unstable.
Model routing is a small wrinkle. Calls requested specific DeepSeek V4 routes but returned wrappers did not record the resolved model ID, so exact serving behavior is logged by request, not confirmation.
Finally, the Flash crossover’s seven card-vs-description contrasts beyond the frozen one are exploratory and unadjusted for multiplicity. Treat those numbers as suggestive, not confirmatory.