Get Started
Home
Topics
Search
Library
6 min read · Agents · Evaluation · Added Oct 5 · Paper published Sep 29, 2026

Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

Source: research paper via Hugging Face Daily Papers
0:00 / 7:51
Do agent-written experiment rationales actually predict outcomes, or just sound good? A paired forecaster test across 336 states found matched explanations gave no confirmed accuracy gain over description-only or donor rationales, though structured prediction cards cut call-to-call drift ~60% — self-consistency without insight.
TL;DR
Predictive Credit asks whether an AI research agent’s written explanation of a planned experiment actually helps forecast that experiment’s outcome, by comparing matched explanations against description-only and donor (someone else’s) explanations under the same forecaster and outcome. Across 336 prospective states, the matched explanation showed no confirmed predictive gain.
Why It Matters
Autonomous research agents now propose experiments, run them, and write up why they expected the result. Benchmarks today score the result. Nobody scores the explanation. If an agent writes “stronger regularization will close the train-test gap because X,” that rationale is either genuinely informative about the outcome or decorative text. Treating all explanations as equally valuable rewards fluent storytelling over actual insight.
The authors want a measurement protocol, not a new agent. Prior simulatability work (notably Hase on Leakage-adjusted simulatability) scored explanations by whether they help an observer predict a model’s output. This paper pushes the same idea onto predicting executed experimental outcomes, which is what research agents are supposed to be good at.
How It Works
The setup is a paired forecast. For each experiment i with public state x, planned intervention a, explanation h, and realized outcome y, a fixed forecaster LLM sees one of three contexts:
•
Description D = (x, a). The baseline: facts and the planned change, no rationale.
•
Matched M = (x, a, h). Adds the agent’s own explanation.
•
Donor S = (x, a, h from another case). Adds somebody else’s explanation, grabbed from a different experiment under a frozen shuffle.
The gain of matched over description measures whether the explanation adds anything at all. The gain of matched over donor tests whether the gain is specific to this case, or whether any plausible-sounding rationale would do. Both use the same outcome and loss, so differences come from the explanation text.
Explanations are elicited as a prediction card with six slots: direction, quantitative target, intermediate observable, benefit regime, falsifier, counterfactual. This structure makes commitments testable. Each forecast returns a point estimate and an 80% interval, scored with Interval score, along with repeat drift between two independent calls.
Five checks are tracked together: commitment (did the card get filled), delivery (did evidence-backed content reach the forecaster), predictive value (D vs M), alignment (S vs M), and sensitivity to known signals.
for state in states: for context in [D, M, S]: y_hat_1, interval_1 = forecaster(context) # two independent calls y_hat_2, interval_2 = forecaster(context) drift[context] = abs(y_hat_1 - y_hat_2) err[context] = mae([y_hat_1, y_hat_2], y_true) iscore[context] = interval_score(interval_1, interval_2, y_true) gain_desc = R(D) - R(M) # matched vs description gain_align = R(S) - R(M) # matched vs donor
Studies run on three settings: a 120-state controlled v5 learning environment, Tox21 with 72 states across 12 assay endpoints, and 24 OpenML tabular tasks giving 144 states. The forecaster is DeepSeek-V4-Pro, with a follow-up DeepSeek-V4-flash replay.
What They Found
The headline is a null, honestly reported. The frozen Tox21 primary test (does the matched card lower ROC AUC Interval score by at least .005) failed: estimate −.0026, 95% interval [−.0174, +.0104]. OpenML’s joint rule on card formation, point-equivalence, and repeatability also failed. v5 was inconclusive. Donor-vs-matched intervals on Tox21 and OpenML spanned zero. So matched explanations did not reliably beat donor explanations on predictive accuracy.
What did move: structured cards sharply cut repeat drift. On Tox21 Pro, drift fell from .01004 (description) to .00357 matched and .00411 donor, reductions of 64.5% and 59.1%. A Flash replay reproduced that drift reduction but raised matched point MAE from .01823 to .02020 and missed interval-score equivalence. Translation: cards make the model self-consistent across calls without making it more accurate.
Elicitation changes commitment rates massively. v5 produced complete six-slot cards in 0/60 spontaneous proposals vs 59/60 when explicitly asked. On OpenML, 78/144 elicited cards came back as empty scaffolds, and assigning the full card widened 80% intervals by 21% while coverage dropped slightly (49.3% vs 51.4% for description).
Direction accuracy mostly came from a sign prior: v5 agents predicted “improvement” in 109/111 cases, scoring 58.6% vs a constant-positive baseline of 57.7%. The forecaster isn’t really discriminating; it’s betting on the prior.
Two positive controls show the forecaster can use strong information when present. A researcher-authored note disclosing the latent functional form in a quadratic-feature experiment cut point MAE by 2.60 pp vs description and 4.36 pp vs a deliberately false mechanism note. A separate Exact-signal calibration where the true outcome was handed in as a hint achieved Spearman .99996. So the pipeline detects real signal; the agent explanations just don’t carry much.
What’s Useful
•
If you’re building a research-agent benchmark, this gives you a protocol to score rationales alongside results. The paired D/M/S design with a fixed forecaster and loss is the reusable piece. You do not need the specific Tox21 or OpenML tasks.
•
Separate repeatability from accuracy in your own evaluations. This paper shows a case where structured prompting cuts call-to-call drift by ~60% while leaving point error flat or slightly worse. Reporting only drift would overclaim.
•
Treat agent-written “direction” claims with suspicion. Compare to a constant-direction baseline before celebrating any direction-accuracy number. The sign prior absorbed most of the signal here.
•
Worth testing in your own setup: elicited cards may improve downstream auditability (explicit falsifiers, counterfactuals) even when they don’t improve forecasts. The paper measures prediction, not usefulness for humans reviewing a proposal.
•
The positive controls suggest a diagnostic: before trusting that an agent’s natural explanation is informative, inject a hand-authored mechanism note on the same tasks and confirm the forecaster responds to it. If your pipeline fails that check, no explanation experiment will work.
Caveats
The null is specific to this forecaster, these tasks, and these donor resolutions (within-cell seed swaps on Tox21 and OpenML, cross-intervention on v5). The paper is careful to call natural-explanation credit “unconfirmed at the tested donor resolutions,” not refuted. A different forecaster or coarser donor pool could change the picture.
Delivery was weak on OpenML: more than half of elicited cards were empty scaffolds after extraction, so the D-vs-M comparison is partly a test of whether cards even arrived with content. The paired-content subset is only 26 states and the authors flag it as unstable.
Model routing is a small wrinkle. Calls requested specific DeepSeek V4 routes but returned wrappers did not record the resolved model ID, so exact serving behavior is logged by request, not confirmation.
Finally, the Flash crossover’s seven card-vs-description contrasts beyond the frozen one are exploratory and unadjusted for multiplicity. Treat those numbers as suggestive, not confirmatory.
Topics
Agents
Evaluation
Agents
Evaluation
Up next in Agents
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents216 episodes
Evaluation153 episodes