An autoresearch loop lets a coding agent rewrite a molecular geometry optimizer against two admission gates, producing AutoSella variants that need 40–77% of the baseline’s force evaluations at DFT level despite never seeing a DFT gradient during the search.
In quantum chemistry, finding a molecule’s relaxed shape means repeatedly asking a physics model for the forces on each atom, then taking a step downhill. At the DFT (Density Functional Theory) level, each force call can dominate wall-clock time, so the number of steps the optimizer needs sets the cost of the whole job. Decades of hand-tuned optimizers (BFGS variants, internal-coordinate tricks, model Hessians) have squeezed this down, and Sella is currently the fastest open-source option the authors found across 27 configurations.
The practical question: can a language model agent, given Sella’s source code and a benchmark, invent a faster optimizer than the humans did? Prior work like FunSearch and AlphaEvolve has shown LLMs can evolve algorithms against a scorer, but applying it to chemistry software raises two obvious failure modes. The agent could “save” force calls by stopping early (worse answers, fewer steps), or it could overfit mechanisms to the training molecules.
The agent gets one editable file (a single-file version of Sella), a fixed evaluator, and a per-molecule results table after every run. Each cycle it reads prior notes, proposes one coherent edit, submits for scoring, and either keeps or discards the change. Two gates do the work:
•
Validity gate: a candidate must reach at least the baseline’s mean energy recovery on training molecules, with no evaluation errors. This blocks the “stop early to save calls” exploit.
•
Generalization gate: if training passes, the candidate is re-run on a held-out validation split. It only replaces the champion if it also wins there. This blocks overfitting to the training set.
Search happens on GFN2-xTB, a cheap semi-empirical potential where a full training-set evaluation takes 60–90 seconds. The expensive DFT evaluations are only used later, to test whether improvements transfer. The agent gets per-molecule diagnostics (which molecules it helped, which it hurt, convergence status), which the authors argue turns the next cycle into a chemistry-grounded hypothesis rather than blind hyperparameter tuning.
champion = vendored_sella_2_5_0
while True:
candidate = agent.edit(champion, log, rejected)
c_train, r_train = evaluate(candidate, train)
if r_train < 1 or has_errors(c_train):
discard(candidate) # validity gate
elif c_train >= champion.cost - 1e-4:
discard(candidate) # no significant gain
else:
c_val, r_val = evaluate(candidate, val)
if r_val >= 1 and c_val < champion.val_cost - 1e-4:
champion = candidate # generalization gate
else:
discard(candidate)
Two independent runs of 66 active research hours each were conducted, one driven by Claude Fable 5.1 (max) via Claude Code and one by GPT-6 Astra (high) via Codex CLI, producing AutoSella-F and AutoSella-A.
AutoSella-F consistently cut force calls. On the GFN2-xTB benchmark it was evolved on, it needed 39.86–73.25% of Sella’s per-molecule force calls across four held-out datasets (drug-like molecules, dipeptides, amino acid–ligand pairs, solvated PubChem), while recovering slightly more energy than the reference on average.
Transfer to unseen potentials held up. Re-running the frozen optimizers on GFN-FF (a classical force field) and r2SCAN-3c DFT, which the agent never evaluated against, the savings persisted. On solvated PubChem under DFT, AutoSella-F needed only 40.24% of Sella’s force calls with 1.27% better mean energy recovery. Because DFT force calls dominate wall time, this reduction translates roughly proportionally to total job cost.
AutoSella-A was weaker and uneven, using 78–93% of baseline calls on three datasets but regressing to 118.77% on solvated PubChem with slightly worse energy recovery. The two agents also worked differently: Fable committed 108 candidates (37 accepted), Astra committed 672 (29 accepted). Fable spent long stretches analyzing saved structures inside its container, while Astra iterated fast.
Ablations identified three load-bearing changes in AutoSella-F: better curvature models and coordinates (removing them adds ~18.7 pp to validation cost), explicit handling of disconnected fragments (~11.3 pp), and smarter reuse of past gradient information in Hessian updates (~10.2 pp). A third run with Grok 4.6 barely beat baseline (97.79% on training) and regressed badly on multi-fragment systems, so the authors treat it as evidence that frontier model choice matters a lot here.
•
If you run DFT geometry optimizations on drug-like molecules, trying AutoSella-F is a plausible cost reduction. The savings were measured on held-out molecules and on a DFT level the search never touched. Multi-fragment and solvated systems benefit most; single-molecule gains are real but smaller.
•
If you’re thinking about autoresearch for your own numerical code, the gated loop pattern is worth studying. The two-gate structure (reject early-stopping hacks, reject overfitting) is a reusable template whenever your objective can be gamed by producing worse outputs faster. The paper’s harness isolates the agent’s workspace from the scorer, which is what keeps the loop honest.
•
Worth testing before adopting: whether AutoSella-F’s gains survive on your chemistry (metal complexes, transition states, periodic systems) are all outside what was evaluated. The paper only tested drug-design-adjacent molecules with the convergence criteria from geomeTRIC.
•
The surrogate-potential trick generalizes. Searching on a cheap potential (GFN2-xTB) and transferring to the expensive one (DFT) worked here. If you have a cheap proxy that preserves the structure of your real objective, you can do a lot of search cycles you otherwise couldn’t afford.
•
The baseline is one specific optimizer (Sella 2.5.0 with a specific coordinate option) at one potential level with one convergence test. The authors show Sella is the fastest in their panel of 27, but a different starting point might have yielded different edits or smaller headroom.
•
AutoSella-A regressed on solvated systems, and only ~1.0–39.8% of its solvated-PubChem endpoints match Sella’s local minimum geometrically. Lower final energy sometimes means a different minimum rather than a deeper version of the same one. The authors run RMSD and vibrational checks and flag this honestly.
•
Agent choice dominates outcome. Fable and Astra produced usefully different optimizers; Grok 4.6 barely moved past baseline on the same harness and budget. Reproducing this with a different frontier model is not guaranteed to work.
•
Compute cost of the search itself (66 active hours per run across 192 distributed workers, plus whatever the LLM API cost) is not amortized against the per-job savings in the paper. Whether this pays off depends on how many DFT optimizations you plan to run with the resulting optimizer.