{"id":"savings-surrogate-neb","title":"Surrogate-Accelerated Transition-State and Pathway Search","subtitle":"GPR-NEB and variants, ML-surrogate NEB, surrogate dimer methods, and uMLIP-warmed practice, 2015–2026; costs counted in electronic-structure evaluations.","category":"references","tags":["literature-review","savings-stack","neb","surrogate","transition-state"],"source":"articles/lit-review/savings-surrogate-neb.md","lang":"en","words":2697,"readMinutes":12,"toc":[{"depth":2,"text":"1. Executive result","id":"1-executive-result"},{"depth":2,"text":"2. Savings table","id":"2-savings-table"},{"depth":2,"text":"3. Proven vs claimed","id":"3-proven-vs-claimed"},{"depth":2,"text":"4. Openings for Lupine","id":"4-openings-for-lupine"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> explore agent <code>agent-13</code> (director-commissioned deep research, swarm of 7, 2026-07-21) — materialized verbatim, then editorially corrected only to replace the retracted 624/132 union-anchor figures with the reproducible primary-record values (558/154). Quantitative literature claims are as reported by the research agent from sources it accessed; see citations inline. Citation-verification pass pending before any external publication.</p>\n</blockquote>\n<h1 id=\"deep-research-digest-surrogate-accelerated-transition-state-and-pathway-search\">Deep-research digest: Surrogate-accelerated transition-state and pathway search</h1><p>Scope covered: GPR-NEB and variants, NN/ML-surrogate NEB, surrogate-assisted dimer / minimum-mode following, and uMLIP-warmed NEB practice, 2015–2026 (evidence cut 2026-07-21). Cost accounting throughout follows the literature convention: <strong>number of energy+force (electronic structure) evaluations to convergence</strong>, since that dominates wall time (&quot;the overall computational effort is well characterized by simply the number of times the energy and force need to be evaluated,&quot; Koistinen et al. 2017). 22 real sources accessed this session; abstract-only vs full-text status is stated per row.</p>\n<h2 id=\"1-executive-result\">1. Executive result</h2><ul>\n<li><strong>Local (per-search) GP surrogates are the only technique with an independently replicated, decade-long track record of ~one-order-of-magnitude reductions in true energy/force evaluations for saddle searches, with no measured loss in converged barrier accuracy.</strong> Koistinen/Jónsson 2017 (JCP 147, 152720) showed an order-of-magnitude cut on the 13-transition heptamer-island benchmark; Garrido Torres et al. 2019 (PRL 122, 156001) showed 5–25× fewer calls vs FIRE/LBFGS/MDMin on three metal-surface transitions (243→11 calls on Müller-Brown) and, crucially, <strong>convergence cost decoupled from the number of band images</strong>; Goswami et al. 2025 (JCTC 21, 7935) repeated the <del>10× result on <strong>500 molecular reactions</strong> for the GP-dimer. The 2026 tutorial review (arXiv:2603.10992) summarizes: saddle searches drop from &quot;hundreds of calls to tens&quot; (</del>30 evaluations typical).</li>\n<li><strong>Pretrained-GNN (&quot;one-shot uMLIP&quot;) NEB gives bigger headline factors but with a real accuracy price.</strong> CatTSunami (arXiv:2405.02078): OC20-trained GNN finds TS energies within 0.1 eV of DFT <strong>91% of the time at 28× speedup</strong> on the 932-calculation OC20NEB benchmark; dense enumeration of a 174-reaction network cost 12 GPU-days vs an estimated 52 GPU-years of DFT (~1500×) — but that figure is throughput amortization, not per-search.</li>\n<li><strong>Raw uMLIP-NEB barrier errors are too large for kinetics without correction.</strong> The largest head-to-head to date (574 migration paths, battery chemistries; Digital Discovery 2026, DOI 10.1039/D5DD00534E): MAE vs DFT-NEB is 0.310 eV (MACE-MP-0) to 0.349 eV (M3GNet); best-case Orb-v3 0.198 eV after outlier exclusion; CHGNet/M3GNet systematically <em>underestimate</em> barriers (73–78% of predictions). Useful for 500-meV-threshold screening (≤85% classification accuracy) and for <strong>warm-starting</strong> DFT-NEB (MACE-MP-0/SevenNet-relaxed paths beat linear interpolation in &gt;71% of cases; fewer ionic+electronic steps in 5 of 6 restarted DFT-NEBs), not for direct barrier values.</li>\n<li><strong>Active-learning MLIP protocols convert a small, fixed DFT budget into DFT-grade NEBs.</strong> Schaaf et al. (arXiv:2301.09931): 622 total DFT single-points train a GAP that reproduces all five CO₂-to-methanol/In₂O₃ barriers within 45 meV, and DFT-NEB restarted from the ML minimum-energy path <strong>converges with zero optimization steps</strong>. Fine-tuned uMLIPs (Lian et al., J. Mater. Chem. A 2025) fix the pretrained-potential bias on high-energy states for battery screening.</li>\n<li>The classical line&#39;s failure mode is known and actively patched: stationary kernels break on short-distance repulsion (fixed by inverse-distance kernels, 2019), and GP variance is a sampling-density signal, <strong>not</strong> an accuracy bound (stated plainly in the 2026 review) — trust regions and pruning (arXiv:2510.06030, mean time halved again on 238 hard reactions) are the current state of the art.</li>\n</ul>\n<h2 id=\"2-savings-table\">2. Savings table</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Technique</th>\n<th>What cost it removes</th>\n<th>Measured savings factor (system/size context)</th>\n<th>Accuracy cost / failure mode</th>\n<th>Citation (access level)</th>\n<th>Evidence strength</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Technique\">GP-NEB, squared-exponential kernel, one-image-evaluated acquisition</td>\n<td data-label=\"What cost it removes\">Iterative force calls on all images during band relaxation</td>\n<td data-label=\"Measured savings factor (system/size context)\">~10× fewer energy+force evaluations vs regular NEB; 13 rearrangement transitions of a heptamer island on FCC(111) (EAM); uncertainty-targeted evaluation halves calls further</td>\n<td data-label=\"Accuracy cost / failure mode\">None measured for converged MEP; stationary kernel fails when training data includes close-contact, large-force configs</td>\n<td data-label=\"Citation (access level)\">Koistinen et al., JCP 147, 152720 (2017), <a href=\"https://arxiv.org/abs/1706.04606\">arXiv:1706.04606v2</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Strong (peer-reviewed; later replicated by other groups)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GP-NEB, inverse-interatomic-distance kernel + early stopping</td>\n<td data-label=\"What cost it removes\">Same; plus failures/restarts on repulsive regions</td>\n<td data-label=\"Measured savings factor (system/size context)\">Fewer evaluations than 2017 version on the same heptamer benchmark; enables cases where 2017 fails: H₂ dissociation on Cu(110) (216-atom slab, EAM), H₂O hop on ice Ih(0001); per-system call counts not in abstract [UNVERIFIED — abstract-level only]</td>\n<td data-label=\"Accuracy cost / failure mode\">Fixes, not introduces, failure modes</td>\n<td data-label=\"Citation (access level)\">Koistinen et al., JCTC 15, 6738 (2019), DOI <a href=\"https://pubs.acs.org/doi/10.1021/acs.jctc.9b00692\">10.1021/acs.jctc.9b00692</a>; <a href=\"https://chemrxiv.org/engage/api-gateway/chemrxiv/assets/orp/resource/item/60c7458d9abda2f58df8c5cb/original/nudged-elastic-band-calculations-accelerated-with-gaussian-process-regression-based-on-inverse-inter-atomic-distances.pdf\">ChemRxiv full text</a> (full text read)</td>\n<td data-label=\"Evidence strength\">Strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ML-NEB (CatLearn GP surrogate, acquisition-function NEB)</td>\n<td data-label=\"What cost it removes\">Cost scaling with number of images; image-count tuning</td>\n<td data-label=\"Measured savings factor (system/size context)\">243→11 force calls on Müller-Brown (≈22×); ~5–25× fewer function calls vs FIRE/LBFGS/MDMin on Au adatom/Al(111), Pt adatom/stepped Pt, Pt heptamer island/Pt(111) (all EMT); &quot;order of magnitude faster… no accuracy loss for converged energy barriers&quot;</td>\n<td data-label=\"Accuracy cost / failure mode\">Demonstrated rigorously only on EMT/toy PES; DFT validation confined to SI; LBFGS baseline occasionally fails</td>\n<td data-label=\"Citation (access level)\">Garrido Torres et al., PRL 122, 156001 (2019), <a href=\"https://arxiv.org/abs/1811.08022\">arXiv:1811.08022</a> / <a href=\"https://ar5iv.labs.arxiv.org/html/1811.08022\">full text</a> (full text read)</td>\n<td data-label=\"Evidence strength\">Strong (toy/EMT); moderate (DFT)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">NN surrogate saddle search (Amp)</td>\n<td data-label=\"What cost it removes\">Intermediate ab initio force calls during saddle search</td>\n<td data-label=\"Measured savings factor (system/size context)\">&quot;Dramatic reduction&quot; on two simple examples; exact factor not in abstract [UNVERIFIED]</td>\n<td data-label=\"Accuracy cost / failure mode\">No in-time validation; uncertainty unavailable (criticized by later work)</td>\n<td data-label=\"Citation (access level)\">Peterson, JCP 145, 074106 (2016), DOI <a href=\"https://pubmed.ncbi.nlm.nih.gov/27544086/\">10.1063/1.4960708</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate (seminal, superseded)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GP minimum-mode following (GPR-accelerated dimer)</td>\n<td data-label=\"What cost it removes\">Rotation-phase force evaluations (5–15 per translation step)</td>\n<td data-label=\"Measured savings factor (system/size context)\">Evaluations &quot;reduced to less than a third&quot; vs conventional; starts near saddles for H₂/Cu(110) dissociative adsorption + 3 gas-phase reactions</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation (access level)\">Koistinen et al., JCTC 16, 499 (2020), DOI <a href=\"https://pubmed.ncbi.nlm.nih.gov/31801018/\">10.1021/acs.jctc.9b01038</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPR TS search (local, molecular)</td>\n<td data-label=\"What cost it removes\">Gradient evaluations vs dimer and P-RFO</td>\n<td data-label=\"Measured savings factor (system/size context)\">&quot;Significant decrease&quot; in energy/gradient evaluations on 27 test systems (exact factor not in abstract)</td>\n<td data-label=\"Accuracy cost / failure mode\">Local method; needs decent starting point (they supply one)</td>\n<td data-label=\"Citation (access level)\">Denzel &amp; Kästner, JCTC 2018, DOI <a href=\"https://api.semanticscholar.org/graph/v1/paper/DOI:10.1021/acs.jctc.8b00708?fields=title,abstract,year,authors\">10.1021/acs.jctc.8b00708</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Physical prior-mean GP + CI-NEB</td>\n<td data-label=\"What cost it removes\">Optimizer steps vs FIRE</td>\n<td data-label=\"Measured savings factor (system/size context)\">~10× fewer optimization steps than FIRE across gas-phase, bulk-phase, and interfacial reaction benchmarks</td>\n<td data-label=\"Accuracy cost / failure mode\">Surrogate PES &quot;high accuracy vs true PES&quot;; single-group result</td>\n<td data-label=\"Citation (access level)\">Teng, Wang, Bao, JCTC 20, 4308 (2024), DOI <a href=\"https://pubmed.ncbi.nlm.nih.gov/38720441/\">10.1021/acs.jctc.4c00291</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GP-dimer production implementation (eOn/gpr_optim, C++)</td>\n<td data-label=\"What cost it removes\">Electronic structure calls in single-ended search</td>\n<td data-label=\"Measured savings factor (system/size context)\"><strong>Order-of-magnitude fewer electronic structure evaluations than dimer across a 500-molecular-reaction benchmark</strong> (Hermez et al. set); Cartesian GP competitive with Sella internal coordinates; wall time reduced in 3/4 cases even at cheap HF level</td>\n<td data-label=\"Accuracy cost / failure mode\">GP hyperparameter overhead can dominate cheap oracles; failures if search strays outside data region</td>\n<td data-label=\"Citation (access level)\">Goswami et al., JCTC 21, 7935 (2025), <a href=\"https://arxiv.org/abs/2505.12519\">arXiv:2505.12519v2</a>, DOI 10.1021/acs.jctc.5c00866 (abstract)</td>\n<td data-label=\"Evidence strength\">Strong (large benchmark)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Adaptive-pruning OT-GP (Wasserstein farthest-point subset, trust radius)</td>\n<td data-label=\"What cost it removes\">GP update overhead and instability as data grows</td>\n<td data-label=\"Measured savings factor (system/size context)\">Mean computational time reduced to less than half on 238 challenging reaction configurations</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation (access level)\">Goswami &amp; Jónsson, ChemPhysChem 2025, <a href=\"https://arxiv.org/abs/2510.06030\">arXiv:2510.06030v3</a>, DOI 10.1002/cphc.202500730 (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate→strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">On-the-fly GPR hybrid calculator (ASE add-on)</td>\n<td data-label=\"What cost it removes\">Fraction of DFT calls inside NEB replaced by surrogate predictions</td>\n<td data-label=\"Measured savings factor (system/size context)\">3–10× acceleration vs pure ab initio NEB on surface diffusion and reactions (per-system breakdown not in abstract)</td>\n<td data-label=\"Accuracy cost / failure mode\">Uncertainty-threshold behavior governs accuracy; one-shot MLFFs criticized as unsuitable for high-accuracy exploratory tasks</td>\n<td data-label=\"Citation (access level)\">Onyango, Kang &amp; Zhu, Comput. Phys. Commun. 2025, <a href=\"https://arxiv.org/abs/2504.07319\">arXiv:2504.07319v2</a> (abstract + methods full text)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPR + LI-NEB instanton path optimization</td>\n<td data-label=\"What cost it removes\">On-the-fly electronic structure calls in ring-polymer instanton search</td>\n<td data-label=\"Measured savings factor (system/size context)\">Order of magnitude faster than traditional instanton algorithms; H + CH₄ → H₂ + CH₃, malonaldehyde, aminopropenal</td>\n<td data-label=\"Accuracy cost / failure mode\">Tunneling rates in &quot;excellent agreement&quot; (no degradation reported)</td>\n<td data-label=\"Citation (access level)\">Zhang et al., JCTC 21, 7517 (2025), DOI <a href=\"https://pubmed.ncbi.nlm.nih.gov/40674652/\">10.1021/acs.jctc.5c00673</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate (single group, new area)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Gradient-enhanced Kriging TS/constrained optimization</td>\n<td data-label=\"What cost it removes\">Energy/gradient evaluations in TS and reaction-path optimization</td>\n<td data-label=\"Measured savings factor (system/size context)\">&quot;Outperforms current standard in efficiency and robustness&quot; on a reaction benchmark set (counts not in abstract [UNVERIFIED])</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation (access level)\">Raggi, Fdez. Galván, Lindh, JCTC 17, 193 (2021), DOI <a href=\"https://api.semanticscholar.org/graph/v1/paper/DOI:10.1021/acs.jctc.0c01163?fields=title,abstract,year,authors\">10.1021/acs.jctc.0c01163</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPR surrogate + gentlest ascent dynamics, OED active learning</td>\n<td data-label=\"What cost it removes\">True-gradient evaluations in single-ended saddle search</td>\n<td data-label=\"Measured savings factor (system/size context)\">&quot;Competitive performance&quot; on three examples; no DFT system, no factor in abstract</td>\n<td data-label=\"Accuracy cost / failure mode\">Math/test-function level</td>\n<td data-label=\"Citation (access level)\">Gu, Wang, Zhou, J. Sci. Comput. 93, 78 (2022), <a href=\"https://ar5iv.labs.arxiv.org/html/2108.04698\">arXiv:2108.04698</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">On-the-fly targeted MLIP inside ART nouveau (single-ended saddle sampling)</td>\n<td data-label=\"What cost it removes\">Reference-potential evaluations during activated-mechanism exploration</td>\n<td data-label=\"Measured savings factor (system/size context)\">Targeted on-the-fly MLIP gives highest barrier precision while &quot;remaining cost-effective&quot; (no factor in abstract); SW Si vacancy + SiGe zincblende</td>\n<td data-label=\"Accuracy cost / failure mode\">General-purpose potentials underperform targeted ones for barriers</td>\n<td data-label=\"Citation (access level)\">Sanscartier et al., JCP 158, 244110 (2023), DOI 10.1063/5.0143211, <a href=\"https://arxiv.org/abs/2301.08630\">arXiv:2301.08630</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">AL-trained GAP MLFF + NEB (system-specific)</td>\n<td data-label=\"What cost it removes\">Nearly all DFT-NEB force calls for a reaction set</td>\n<td data-label=\"Measured savings factor (system/size context)\"><strong>622 total DFT single-points</strong> yield all five CO₂→methanol/In₂O₃ barriers within 45 meV (after 6 NEB-AL iterations); DFT-NEB restarted from ML MEP converges in <strong>0 optimization steps</strong> (projected force &lt; 0.05 eV/Å); found 40%-lower rate-limiting barrier than literature</td>\n<td data-label=\"Accuracy cost / failure mode\">Pre-AL MLFF: mean barrier error 50 meV; protocol-specific, needs MD+AL pipeline</td>\n<td data-label=\"Citation (access level)\">Schaaf et al. 2023, <a href=\"https://arxiv.org/abs/2301.09931\">arXiv:2301.09931v2</a> (abstract + full-text excerpts)</td>\n<td data-label=\"Evidence strength\">Strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Pretrained GNN (OC20) as NEB calculator (CatTSunami)</td>\n<td data-label=\"What cost it removes\">All per-image DFT calls in catalytic NEB</td>\n<td data-label=\"Measured savings factor (system/size context)\">28× speedup; TS within 0.1 eV of DFT in 91% of OC20NEB cases (932 DFT NEBs); 174-reaction network at 40 meV resolution: 12 GPU-days vs <del>52 GPU-years (</del>1500×)</td>\n<td data-label=\"Accuracy cost / failure mode\">9% of cases miss 0.1 eV; no explicit reaction training; fixed to OC20 DFT settings</td>\n<td data-label=\"Citation (access level)\">Wander et al., <a href=\"https://arxiv.org/abs/2405.02078\">arXiv:2405.02078v3</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Strong (large public benchmark; journal version exists, venue detail [UNVERIFIED])</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GNN surrogate NEB for molecules (PaiNN/Transition1x)</td>\n<td data-label=\"What cost it removes\">DFT calls in organic-reaction NEB</td>\n<td data-label=\"Measured savings factor (system/size context)\">Barrier MAE 0.13 ± 0.03 eV on unseen reactions; outperforms DFTB on accuracy and cost</td>\n<td data-label=\"Accuracy cost / failure mode\">QM9/ANI1x-trained models fail to converge (data in TS region is the binding constraint)</td>\n<td data-label=\"Citation (access level)\">Schreiner et al., <a href=\"https://arxiv.org/abs/2207.09971\">arXiv:2207.09971v3</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">uMLIP-NEB as-is (MACE-MP-0, MACE-OMAT, Orb-v3, SevenNet, CHGNet, M3GNet)</td>\n<td data-label=\"What cost it removes\">All DFT-NEB calls, if trusted</td>\n<td data-label=\"Measured savings factor (system/size context)\">574 literature migration paths (battery materials): full-replacement at MAE 0.310–0.349 eV; classification (good/bad conductor, 500 meV) ≤84.8% (Orb-v3); warm-start: better-than-LI initial paths in &gt;66% (MACE/SevenNet &gt;71%), fewer DFT steps in 5/6 restarts</td>\n<td data-label=\"Accuracy cost / failure mode\">0.2–0.35 eV MAE ≈ 3–6× the intrinsic DFT-NEB error (~60 meV ≈ 10× in diffusivity at 298 K); CHGNet/M3GNet underestimate 73–78%; all models degrade at high barriers (≤21% within 0.1 eV above ~1.3 eV); geometry and barrier accuracy uncorrelated</td>\n<td data-label=\"Citation (access level)\">Bheemaguli, Xiao, Gautam, Digital Discovery 2026, DOI <a href=\"https://pubs.rsc.org/en/content/articlehtml/2026/dd/d5dd00534e\">10.1039/D5DD00534E</a>, <a href=\"https://arxiv.org/abs/2512.03642\">arXiv:2512.03642</a> (full text read)</td>\n<td data-label=\"Evidence strength\">Strong (largest uMLIP-NEB benchmark)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Fine-tuned uMLIP NEB/MD (dual framework)</td>\n<td data-label=\"What cost it removes\">DFT calls in high-throughput Li-ion screening</td>\n<td data-label=\"Measured savings factor (system/size context)\">High-throughput barrier prediction &quot;balancing speed and accuracy,&quot; validated on NASICON-type Li electrolytes; no call counts in abstract</td>\n<td data-label=\"Accuracy cost / failure mode\">Pretrained uMLIPs explicitly inaccurate on high-energy states → fine-tuning required</td>\n<td data-label=\"Citation (access level)\">Lian et al., J. Mater. Chem. A 13, 34918 (2025), DOI 10.1039/D5TA05355B, <a href=\"https://arxiv.org/abs/2507.02334\">arXiv:2507.02334</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">uMLIP grid-PES sampling, no NEB images (FastTrack)</td>\n<td data-label=\"What cost it removes\">NEB machinery and all per-image DFT calls</td>\n<td data-label=\"Measured savings factor (system/size context)\">~10²× speedup over DFT-NEB; barriers within &quot;tens of meV&quot; of DFT/experiment; 12 electrode/electrolyte materials (LiCoO₂, LiFePO₄, LGPS); GPTFF/CHGNet/MACE compared</td>\n<td data-label=\"Accuracy cost / failure mode\">Preprint; single paper; confined to single-ion hops in crystals</td>\n<td data-label=\"Citation (access level)\">Liu et al., <a href=\"https://arxiv.org/abs/2508.10505\">arXiv:2508.10505</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">AL + CI-NEB slip-pathway workflow (interlayer slip PES)</td>\n<td data-label=\"What cost it removes\">Dense DFT sampling of slip PES</td>\n<td data-label=\"Measured savings factor (system/size context)\">Automated slip-pathway identification across several ductile inorganic semiconductors; no factor in abstract</td>\n<td data-label=\"Accuracy cost / failure mode\">Workflow-level, not per-call accounting</td>\n<td data-label=\"Citation (access level)\">Luo et al., npj Comput. Mater. 11, 41 (2025), DOI <a href=\"https://www.nature.com/articles/s41524-025-01531-7\">10.1038/s41524-025-01531-7</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Agentic DFT TS workflow (LLM-driven, UMA-warmed)</td>\n<td data-label=\"What cost it removes\">Human-in-the-loop cost; some wasted DFT via replanning</td>\n<td data-label=\"Measured savings factor (system/size context)\">83% TS success on 100-reaction OC20NEB subset; vs human experts 70% vs 73±12% on 10 held-out; baseline ~10,000 CPU-hours per DFT TS search</td>\n<td data-label=\"Accuracy cost / failure mode\">Not a call-count method; shows current DFT cost baseline</td>\n<td data-label=\"Citation (access level)\">Madhavan et al., <a href=\"https://arxiv.org/abs/2605.14154\">arXiv:2605.14154v1</a> (abstract)</td>\n<td data-label=\"Evidence strength\">Claim (preprint)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Unified BO/GP framing of minimization–dimer–NEB (review)</td>\n<td data-label=\"What cost it removes\">—</td>\n<td data-label=\"Measured savings factor (system/size context)\">Consensus estimate: &quot;roughly an order of magnitude&quot; reduction in electronic structure evaluations; ~30 evaluations per saddle search typical with analytical forces; gains shrink for cheap oracles / short searches / energy-only data</td>\n<td data-label=\"Accuracy cost / failure mode\">Variance ≠ accuracy; trust regions required; GP overhead O(M³), M~30</td>\n<td data-label=\"Citation (access level)\">Goswami, <a href=\"https://arxiv.org/html/2603.10992v5\">arXiv:2603.10992v5</a> (2026; full text read)</td>\n<td data-label=\"Evidence strength\">Strong (review of the field)</td>\n</tr>\n</tbody></table></div><p>Sighted but not verified (citation only, no abstract read): Fang, Zhu, Cheng, Hao, Richardson, &quot;Robust Gaussian Process Regression Method for Efficient Tunneling Pathway Optimization: Application to Surface Processes,&quot; JCTC 20, 3766 (2024) — relevant GPR-pathway work from a Chinese group (USTC) + ETH; numbers [UNVERIFIED]. Denzel et al., JPCA 2019 (GPR for MEP optimization + TS search, DOI 10.1021/acs.jpca.9b08239) — fetch failed [UNVERIFIED].</p>\n<h2 id=\"3-proven-vs-claimed\">3. Proven vs claimed</h2><p><strong>Replicated and trusted (multiple independent groups, peer-reviewed, large benchmarks):</strong></p>\n<ul>\n<li>The ~10× evaluation reduction from per-search GP surrogates on both double-ended (NEB) and single-ended (dimer/minimum-mode) searches. Line: Iceland/Aalto 2017/2019/2020 → Stanford/SUNCAT 2019 → Northeastern (Bao) 2024 → Iceland/EPFL (Goswami/Jónsson) 2025 on 500 reactions, with an independent US implementation (UNCC, GPR_calculator, 3–10× on true DFT) and a 2026 review codifying it. Accuracy cost at convergence: none measured, provided the inverse-distance kernel and trust-region/early-stopping safeguards are used.</li>\n<li>Image-count-independent convergence (Garrido Torres): cost of converging the band no longer scales with the number of moving images — a structural, not incremental, saving.</li>\n<li>uMLIP-NEB as a <em>screening/warm-start</em> tool, not a barrier tool: the 574-path Digital Discovery benchmark quantifies the bias (0.2–0.35 eV MAE, systematic underestimation by invariant models) and the warm-start benefit (&gt;71% better initial paths, fewer DFT steps in 5/6 restarts).</li>\n</ul>\n<p><strong>Single-paper claims, plausible but not independently replicated:</strong></p>\n<ul>\n<li>CatTSunami&#39;s 28×/91%-within-0.1 eV (one group, but public OC20NEB dataset now used by others — UMA, TSAgent — so the <em>benchmark</em> is trusted; the 28× figure is theirs). The 1500× is throughput amortization across a huge parallel enumeration — real but not per-search, and should never be quoted as a per-NEB speedup.</li>\n<li>Teng–Bao prior-mean GP ~10× vs FIRE; Zhang–Govind instanton GPR ~10×; Schaaf 622-single-point protocol (a second group would need to rerun the protocol to confirm).</li>\n<li>Sanscartier/Mousseau on-the-fly ARTn MLIP (cost-effectiveness asserted, not quantified in abstract).</li>\n</ul>\n<p><strong>Marketing or thin:</strong></p>\n<ul>\n<li>FastTrack&#39;s &quot;~10²×, tens of meV&quot; (preprint, one group).</li>\n<li>TSAgent&#39;s agentic success rates (preprint; its real value is the honest baseline: a no-surrogate DFT TS search on OC20NEB-like systems costs ~10⁴ CPU-hours).</li>\n<li>Any claim of uMLIP-NEB <em>replacing</em> DFT-NEB for barriers — contradicted by the strongest benchmark available (D5DD00534E).</li>\n<li>Note the honest scaling caveat from the 2026 review: gains depend on oracle cost dominating, search distance being long, and analytical forces being available (3N+1 data per call); minimization and short searches gain far less. GPR overhead (hyperparameter re-optimization) eats the savings when the oracle is cheap — Goswami&#39;s 500-reaction study only won wall time in 3 of 4 cases even at HF cost.</li>\n</ul>\n<h2 id=\"4-openings-for-lupine\">4. Openings for Lupine</h2><ol>\n<li><strong>Theorem-gated anchors can replace the weakest part of the GP-NEB acquisition loop: the uncertainty signal.</strong> The 2026 review admits GP variance is only a sampling-density proxy, not an accuracy bound, and the field patches this with hand-tuned trust radii and early-stopping rules. Lupine&#39;s physical-law theorems (formalized in Lean) are exactly the missing <em>correctness</em> gate: they can certify when a surrogate/uMLIP prediction near a saddle violates a physical constraint (energy ordering, curvature sign, Hessian mode structure), deciding the next DFT call on physics rather than kernel distance. This directly addresses the failure mode documented for raw uMLIP-NEB (0.2–0.35 eV bias, systematic underestimation) without retraining anything.</li>\n<li><strong>Sparse-anchor DFT (~4–6 images near the predicted saddle) is the minimal-cost fix for the exact regime where uMLIP-NEB fails.</strong> Digital Discovery 2026 shows all six uMLIPs share 17 common &gt;1 eV outliers and degrade above ~1.3 eV barriers. A Lupine layer on top of MACE-MP-0/Orb-v3 NEB — run the cheap band, theorem-gate the saddle region, evaluate DFT only at 4–6 anchor images near the model-predicted maximum — should convert the uMLIP&#39;s 0.2–0.35 eV MAE into DFT-grade barriers at a small fraction of the ~10² evaluations/image × 10² steps of CI-NEB. This is a cleaner, quantified value proposition than &quot;another surrogate,&quot; because the failure set is now public and enumerated (github.com/sai-mat-group/mlips-migration-barriers, 574 paths).</li>\n<li><strong>Union anchors attack the one cost the entire surrogate literature cannot: per-search surrogate retraining.</strong> Every GP-NEB/GP-dimer run builds its surrogate from zero data and discards it (&quot;the surrogate is discarded after each search completes,&quot; arXiv:2603.10992). In panel studies (reaction networks, screening campaigns, multi-model comparisons), Lupine&#39;s 154-shared-vs-558-naive-anchor result (72.4% fewer DFT calls) compounds multiplicatively with the per-search ~10×: different uMLIPs warm-start different bands, and wherever their predicted extrema coincide, one theorem-gated DFT anchor validates all of them. Nobody in this literature shares anchor data across searches or across models — Schaaf&#39;s protocol comes closest (622 points amortized over 5 reactions) but is single-model and system-specific.</li>\n<li><strong>Complement, don&#39;t fight, the GP-dimer line for single-ended search.</strong> Goswami&#39;s ~30-evaluations-per-saddle on a corrected oracle is already near the floor for single searches; Lupine&#39;s edge there is to run the dimer/minimum-mode walk on the uMLIP and spend theorem-gated DFT only at the converged saddle (plus a curvature-sign check), rather than correcting every surrogate step. Where Lupine should <em>not</em> compete: single isolated reactions with cheap oracles — local GP surrogates win outright with zero pretraining, and the honest literature says so. Lupine&#39;s wedge is amortized, multi-path, multi-model workloads — precisely the regime (CatTSunami-scale networks, battery screening panels) where per-search surrogates restart from scratch and pretrained-model bias is fixed and uncorrected.</li>\n</ol>\n"}