{"id":"savings-abstention-economics","title":"Uncertainty-Gated Selective DFT and the Economics of Abstention","subtitle":"Measured DFT-call budgets, labeling and abstention fractions, and accuracy-per-call pricing.","category":"references","tags":["literature-review","savings-stack","abstention","uncertainty"],"source":"articles/lit-review/savings-abstention-economics.md","lang":"en","words":2387,"readMinutes":11,"toc":[{"depth":2,"text":"1. Executive result","id":"1-executive-result"},{"depth":2,"text":"2. Savings table","id":"2-savings-table"},{"depth":2,"text":"3. Proven vs claimed","id":"3-proven-vs-claimed"},{"depth":2,"text":"4. Openings for Lupine","id":"4-openings-for-lupine"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> explore agent <code>agent-16</code> (director-commissioned deep research, swarm of 7, 2026-07-21) — materialized verbatim, then editorially corrected only to replace the retracted 624/132 union-anchor figures with the reproducible primary-record values (558/154). Quantitative literature claims are as reported by the research agent from sources it accessed; see citations inline. Citation-verification pass pending before any external publication.</p>\n</blockquote>\n<h1 id=\"digest-uncertainty-gated-selective-dft-and-the-economics-of-abstention\">Digest: Uncertainty-gated selective DFT and the economics of abstention</h1><p>Scope note: per instructions, I did <strong>not</strong> re-survey the certified-inference per-step gating literature (coverage/selectivity theory, gate runtime costs). Everything below is economics: measured DFT-call budgets, labeling/abstention fractions, and accuracy-per-call pricing. Access level per source is marked <strong>FT</strong> (full text read this session) or <strong>ABS</strong> (abstract only). Evidence cut: 2026-07-21.</p>\n<h2 id=\"1-executive-result\">1. Executive result</h2><ul>\n<li><strong>The biggest, most replicated saving is abstention during training data generation, not during inference.</strong> Bayesian-error and ensemble-disagreement gates routinely skip 99%+ of first-principles evaluations: VASP&#39;s on-the-fly MLFF skips &gt;99% of FP steps during force-field generation (Jinnouchi/Kresse, PRL 2019 <strong>FT</strong>; PRB 2019 <strong>ABS</strong>), and DP-GEN labeled only 0.0044% of ~650 million explored Al/Mg/Al-Mg configurations (Zhang/Car/E, PRM 2019 <strong>FT</strong>). Final models retain near-FP accuracy (2.6 meV/atom, 0.07 eV/Å in the perovskite case).</li>\n<li><strong>Query-by-committee gates price data at roughly 10%.</strong> Across two independent ecosystems (LANL&#39;s ANI and Princeton/IAPCM&#39;s DP), ensemble-disagreement selection reaches matched or better accuracy with <del>10% of the data of naive sampling: ANI-1x matched ANI-1 on COMP6 with 10% of the data (Smith et al., JCP 2018 <strong>ABS</strong>), and ANI-1ccx used ~480k QbC-selected CCSD(T) points from a 5M pool (</del>10%) to reach coupled-cluster-level accuracy (Nat. Commun. 2019 <strong>FT</strong> of preprint).</li>\n<li><strong>Reactive-force-field budgets are now countable in hundreds of calls.</strong> FLARE&#39;s SGP-uncertainty gate trained a full reactive H/Pt(111) force field with 250 DFT calls (216 in the reactive run) in ~3 days, then ran 500 ps of two-phase reactive MD reproducing the experimental activation energy (0.25(2) vs 0.21–0.24 eV) (Vandermause et al., npj Comput. Mater. 2022 <strong>FT</strong>). The earlier FLARE paper reports &lt;100 DFT calls per 10 ps on-the-fly run (npj Comput. Mater. 2020 <strong>ABS+snippet</strong>).</li>\n<li><strong>Uncertainty-gated saddle/path searches save an order of magnitude in oracle evaluations.</strong> GP-regression NEB cuts energy/force evaluations by ~10x (Koistinen/Jónsson, JCP 2017 <strong>ABS</strong>), replicated by an independent 2025 on-the-fly GPR surrogate reporting 3–10x NEB acceleration (GPR_calculator, arXiv:2504.07319 <strong>ABS</strong>). This is the closest published analogue to Lupine&#39;s sparse-anchor NEB economics.</li>\n<li><strong>In production screening, the abstention rate is measured and it is ~13%.</strong> AdsorbML&#39;s ML-rank-then-DFT-verify protocol finds the DFT-level adsorption minimum in 87.4% of <del>1000 OC20-Dense systems while doing only top-k DFT single-points (</del>2000x effective speedup; the ~13% &quot;failure&quot; tail is what you still pay full DFT for) (Lan et al., npj Comput. Mater. 2023 <strong>ABS + v3 text</strong>). This is the cleanest published accuracy-per-DFT-call tradeoff curve in the subdomain.</li>\n</ul>\n<h2 id=\"2-savings-table\">2. Savings table</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Technique</th>\n<th>What cost it removes</th>\n<th>Measured savings factor (system/size context)</th>\n<th>Accuracy cost / failure mode</th>\n<th>Citation (access)</th>\n<th>Evidence strength</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Technique\">Bayesian-error gate, on-the-fly MLFF (VASP)</td>\n<td data-label=\"What cost it removes\">FP steps during FF training &amp; generation</td>\n<td data-label=\"Measured savings factor (system/size context)\">99% of FP calculations skipped during training, ~100x time cut even <em>while learning</em>; &gt;99% bypassed for melting-point FFs of Al, Si, Ge, Sn, MgO</td>\n<td data-label=\"Accuracy cost / failure mode\">Force/energy error 2.6 meV/atom, 0.07 eV/Å vs FP; melting points reproduced quantitatively</td>\n<td data-label=\"Citation (access)\">Jinnouchi, Karsai &amp; Kresse, PRB 100, 014105 (2019), arXiv:1904.12961 <strong>ABS</strong>. Related earlier perovskite demonstration: Jinnouchi et al., PRL 122, 225701 (2019) <strong>FT</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Model-deviation (max force σ over 4-model ensemble) gate, DP-GEN</td>\n<td data-label=\"What cost it removes\">DFT labeling of explored configurations</td>\n<td data-label=\"Measured savings factor (system/size context)\">0.0044% of ~650M explored configs labeled (Al, Mg, Al-Mg, 50–2000 K)</td>\n<td data-label=\"Accuracy cost / failure mode\">Uniform accuracy on defects/surfaces/phonons not in training data; failure mode = unlucky ensemble agreement in extrapolation</td>\n<td data-label=\"Citation (access)\">Zhang, Lin, Wang, Car, E, PRM 3, 023804 (2019), arXiv:1810.11890v2 <strong>FT</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">QbC ensemble disagreement, chemical-space AL (ANI)</td>\n<td data-label=\"What cost it removes\">DFT data needed for a general organic-molecule potential</td>\n<td data-label=\"Measured savings factor (system/size context)\">Matches prior potential with 10% of the data; beats it with 25% (COMP6 benchmark, CHNO)</td>\n<td data-label=\"Accuracy cost / failure mode\">Residual bias toward sampled regions; committee σ imperfectly correlated with error</td>\n<td data-label=\"Citation (access)\">Smith et al., JCP 148, 241733 (2018), arXiv:1801.09319 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">QbC + transfer learning to gold-standard QM (ANI-1ccx)</td>\n<td data-label=\"What cost it removes\">CCSD(T)/CBS evaluations</td>\n<td data-label=\"Measured savings factor (system/size context)\"><del>480k coupled-cluster points QbC-selected from a 5M-conformation pool (</del>10%), from 200k seed; +200k DFT via 20 AL iterations on torsions</td>\n<td data-label=\"Accuracy cost / failure mode\">Approaches CCSD(T)/CBS on reaction thermochemistry/torsions; data-hungry in absolute terms</td>\n<td data-label=\"Citation (access)\">Smith et al., Nat. Commun. 10, 2903 (2019) <strong>FT</strong> (ChemRxiv preprint)</td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">SGP local-energy uncertainty gate (FLARE)</td>\n<td data-label=\"What cost it removes\">DFT calls during on-the-fly MD training</td>\n<td data-label=\"Measured savings factor (system/size context)\">&lt;100 DFT calls per 10 ps run across single- and multi-element systems</td>\n<td data-label=\"Accuracy cost / failure mode\">GP regression cost scales with sparse set; error controlled to ~meV/atom RMSE</td>\n<td data-label=\"Citation (access)\">Vandermause et al., npj Comput. Mater. 6, 20 (2020), arXiv:1904.02042v3 <strong>ABS + FT snippet</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">SGP uncertainty gate, reactive FF (H/Pt)</td>\n<td data-label=\"What cost it removes\">DFT calls + human months for reactive FF construction</td>\n<td data-label=\"Measured savings factor (system/size context)\">250 DFT calls total (216 in 3.7 ps reactive run); 65.5 h wall; then 500 ps production MD; &gt;10x post-mapping speedup, 2x faster than ReaxFF</td>\n<td data-label=\"Accuracy cost / failure mode\">Activation energy 0.25(2) eV vs exp 0.21–0.24; C44 within 6% (ReaxFF off by 200%)</td>\n<td data-label=\"Citation (access)\">Vandermause et al., npj Comput. Mater. (2022), PMC9440250 <strong>FT</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Uncertainty-biased MD sampler (UDD-AL)</td>\n<td data-label=\"What cost it removes\">MD steps (hence labels) to reach high-uncertainty/reactive regions</td>\n<td data-label=\"Measured savings factor (system/size context)\">350 K biased run matches 600 K unbiased coverage with far fewer steps; 90 proton transfers/0.5 ns at 350 K vs 0 unbiased (48 at 620 K); glycine dataset built with 1,405 DFT labels, RMSE &lt;0.3 kcal/mol vs 50k-structure reference</td>\n<td data-label=\"Accuracy cost / failure mode\">Two empirical bias parameters (A, B); not automated</td>\n<td data-label=\"Citation (access)\">Kulichenko et al., Nat. Comput. Sci. 3, 230–239 (2023), PMC10766548 <strong>FT</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Uncertainty-biased HAL sampler (ACE)</td>\n<td data-label=\"What cost it removes\">Database assembly cost for alloy/polymer potentials</td>\n<td data-label=\"Measured savings factor (system/size context)\">ACE potential for AlSi10 from 88 DFT configurations (32 atoms each), starting from ~dozen; melting T and density near experiment</td>\n<td data-label=\"Accuracy cost / failure mode\">Bias strength τ heuristic; small-cell bias</td>\n<td data-label=\"Citation (access)\">van der Oord et al., npj Comput. Mater. 9, 168 (2023), arXiv:2210.04225 <strong>ABS + FT snippet</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Extrapolation-grade (maxvol) gate, MTP for alloy convex hulls</td>\n<td data-label=\"What cost it removes\">DFT relaxations in high-throughput screening</td>\n<td data-label=\"Measured savings factor (system/size context)\">3–4 orders of magnitude faster than conventional HT-DFT; found unreported stable structures vs AFLOW (Cu-Pd, Co-Nb-V, Al-Ni-Ti)</td>\n<td data-label=\"Accuracy cost / failure mode\">Factor folds in surrogate speed, not only call reduction; grade misses smooth-interpolation errors</td>\n<td data-label=\"Citation (access)\">Gubaev et al., Comput. Mater. Sci. 156, 148 (2019), arXiv:1806.10567v1 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Δ-ML correction on cheap baseline</td>\n<td data-label=\"What cost it removes\">High-level QM labels for large molecular sets</td>\n<td data-label=\"Measured savings factor (system/size context)\">Semiempirical+ML trained on 1–10% of 134k molecules reproduces DFT-level enthalpies of the remainder (16k C₇H₁₀O₂ isomers to chemical accuracy)</td>\n<td data-label=\"Accuracy cost / failure mode\">Only static properties; baseline must capture physics</td>\n<td data-label=\"Citation (access)\">Ramakrishnan et al., JCTC (2015), arXiv:1503.04987v1 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Δ-ML + rank-compression gate for beyond-DFT MLFF (zirconia)</td>\n<td data-label=\"What cost it removes\">RPA/hybrid-level calculations</td>\n<td data-label=\"Measured savings factor (system/size context)\">168 RPA structures (24-atom cells) selected from 1,275; &lt;150,000 CPU hours total RPA training; could halve further</td>\n<td data-label=\"Accuracy cost / failure mode\">Phase-transition T within ~50 K of experiment; needs DFT-level on-the-fly pass first (592 structures, 96-atom)</td>\n<td data-label=\"Citation (access)\">Liu, Verdi, Karsai, Kresse, PRB 105, L060102 (2022) <strong>FT</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GP-surrogate gate in NEB (GP-NEB)</td>\n<td data-label=\"What cost it removes\">Ab initio energy/force evaluations per MEP</td>\n<td data-label=\"Measured savings factor (system/size context)\">Order-of-magnitude fewer evaluations vs conventional NEB (benchmark: island rearrangement transitions)</td>\n<td data-label=\"Accuracy cost / failure mode\">GP struggles with short-distance large-force configs (fixed in 2019 inverse-distance revision, JCTC 15, 6738; dimer variant JCTC 16, 499 (2020))</td>\n<td data-label=\"Citation (access)\">Koistinen et al., JCP 147, 152720 (2017), arXiv:1706.04606 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">On-the-fly GPR surrogate inside NEB (GPR_calculator)</td>\n<td data-label=\"What cost it removes\">Electronic-structure calls in massive NEB campaigns</td>\n<td data-label=\"Measured savings factor (system/size context)\">3–10x acceleration vs pure ab initio NEB (surface diffusion/reaction demos)</td>\n<td data-label=\"Accuracy cost / failure mode\">Preprint, small demo set; no failure-rate audit</td>\n<td data-label=\"Citation (access)\">Zhu et al., arXiv:2504.07319v2 (2025) <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ML-rank + top-k DFT verify (AdsorbML)</td>\n<td data-label=\"What cost it removes\">DFT relaxations in adsorption-energy screening</td>\n<td data-label=\"Measured savings factor (system/size context)\">87.36% success at ~2000x speedup; balanced option 86.33% at 1331x (OC20-Dense, ~1000 surfaces, ~100k configs)</td>\n<td data-label=\"Accuracy cost / failure mode\">~13% of systems fail to recover the DFT minimum within 0.1 eV; the verify step is the residual DFT cost</td>\n<td data-label=\"Citation (access)\">Lan et al., npj Comput. Mater. 9, 172 (2023), arXiv:2211.16486v3 <strong>ABS + FT snippet</strong></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Ensemble + optimization acquisition to steer DFT (Tran &amp; Ulissi)</td>\n<td data-label=\"What cost it removes\">DFT evaluations in catalyst discovery</td>\n<td data-label=\"Measured savings factor (system/size context)\">Screened alloys of 31 elements (50% of d-block, 33% of p-block); 131 CO₂R candidate surfaces/54 alloys; 258 HER surfaces/102 alloys</td>\n<td data-label=\"Accuracy cost / failure mode\">Exact call-reduction factor <strong>[UNVERIFIED]</strong> — full text paywalled; abstract gives scope, not the factor</td>\n<td data-label=\"Citation (access)\">Tran &amp; Ulissi, Nat. Catal. 1, 696–703 (2018) <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Conformal-calibrated latent-distance UQ (CP+latent)</td>\n<td data-label=\"What cost it removes\">Re-training cost of ensemble/dropout UQ; false alarms</td>\n<td data-label=\"Measured savings factor (system/size context)\">Calibrated <em>and</em> sharp intervals; applied to 1M-image training set at low cost (OC20-family NNFFs)</td>\n<td data-label=\"Accuracy cost / failure mode\">Guarantee void under i.i.d. violation (their own tests); <strong>no DFT-call budget reported</strong> — savings implied, not priced</td>\n<td data-label=\"Citation (access)\">Hu, Musielewicz, Ulissi, Medford, MLST 3, 045028 (2022), arXiv:2208.08337v2 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Statistical error-cutoff for gate thresholds</td>\n<td data-label=\"What cost it removes\">Oracle-error information in threshold setting; cost of ensemble UQ</td>\n<td data-label=\"Measured savings factor (system/size context)\">Cutoff-gated AL with cheap UQ (sparse GP, latent distance) matches true-error-oracle AL on 2 datasets × 3 UQ methods</td>\n<td data-label=\"Accuracy cost / failure mode\">Documents that raw ensemble σ decorrelates from true error — gate on cutoff, not σ</td>\n<td data-label=\"Citation (access)\">Annevelink &amp; Viswanathan, arXiv:2308.15653v1 (2023) <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Calibrated gradient-based (single-model) UQ + uncertainty-biased MD (UBMD)</td>\n<td data-label=\"What cost it removes\">Ensemble compute in AL sampling</td>\n<td data-label=\"Measured savings factor (system/size context)\">Ensemble-comparable or better MLIPs at lower computational cost (alanine dipeptide, MIL-53(Al))</td>\n<td data-label=\"Accuracy cost / failure mode\">Calibration needed; bias parameters manual</td>\n<td data-label=\"Citation (access)\">Zaverkin et al., npj Comput. Mater. 10, 83 (2024), arXiv:2312.01416v2 <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Evidential single-pass UQ gate (eIP / e²IP)</td>\n<td data-label=\"What cost it removes\">Ensemble forward passes in AL/UDD and production monitoring</td>\n<td data-label=\"Measured savings factor (system/size context)\">Reliable UQ &quot;without significant computational overhead&quot;; AL + UDD demos; universal potential with real-time UQ (eIP); equivariant force-covariance version claims better accuracy-efficiency-reliability balance than ensembles (e²IP)</td>\n<td data-label=\"Accuracy cost / failure mode\">Young, single-lineage claims; no independent replication of gate-selectivity numbers</td>\n<td data-label=\"Citation (access)\">Xu et al., arXiv:2407.13994v2 (2024) <strong>ABS</strong>; Wang et al., arXiv:2602.10419v2 (2026) <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">CP-inspired learnable quantile calibration on a foundation model</td>\n<td data-label=\"What cost it removes\">Sharpening of uMLIP intervals before gating</td>\n<td data-label=\"Measured savings factor (system/size context)\">Order-of-magnitude better uncertainty-error correlation; improved AL data efficiency; cross-functional transfer of calibration (MACE-MP-0)</td>\n<td data-label=\"Accuracy cost / failure mode\">Preprint; &quot;negligible cost&quot; unquantified in DFT-call terms</td>\n<td data-label=\"Citation (access)\">Ho et al., arXiv:2510.00721v1 (2025) <strong>ABS</strong></td>\n<td data-label=\"Evidence strength\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Distance-to-training-set adaptive gate in AIMD (Botu &amp; Ramprasad)</td>\n<td data-label=\"What cost it removes\">FP steps in AIMD</td>\n<td data-label=\"Measured savings factor (system/size context)\">Adaptive on-the-fly learning of forces; no factor in abstract <strong>[UNVERIFIED factor]</strong></td>\n<td data-label=\"Accuracy cost / failure mode\">Early descriptor/kernels; superseded by GP/NN gates</td>\n<td data-label=\"Citation (access)\">Botu &amp; Ramprasad, IJQC 115, 1074 (2015) <strong>ABS snippet</strong></td>\n<td data-label=\"Evidence strength\">claim (historical)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Two-member committee NNP gate (Schran/Marsalek)</td>\n<td data-label=\"What cost it removes\">Ensemble size for disagreement UQ; generalization-error control</td>\n<td data-label=\"Measured savings factor (system/size context)\">Committee of 2 NNPs bounds generalization error and drives AL (bulk water, aqueous systems)</td>\n<td data-label=\"Accuracy cost / failure mode\">Number of labeled configs not abstract-quantified <strong>[UNVERIFIED budget]</strong></td>\n<td data-label=\"Citation (access)\">Schran, Brezina, Marsalek, JCP 153, 104105 (2020) <strong>citation + secondary</strong></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n</tbody></table></div><h2 id=\"3-proven-vs-claimed\">3. Proven vs claimed</h2><p><strong>Replicated and trusted (multiple independent groups, numbers reproduced):</strong></p>\n<ul>\n<li><em>Ensemble-disagreement (QbC / model-deviation) gating for training-data economics.</em> Independently built and validated by: LANL (ANI-1x/1ccx, 10%-of-data pricing), Princeton–IAPCM Beijing (DP-GEN, 0.0044% labeling), Harvard MIR (FLARE, hundreds of calls per reactive FF), Cambridge/Ortner-Csányi (HAL, 88-config alloy ACE), Skoltech (MTP maxvol grades, alloy hulls), and Charles Univ. (committee NNPs, Schran 2020). The &quot;~10% of naive data&quot; and &quot;&gt;99% of explored frames skipped&quot; figures are consistent across these — this is the closest thing to a law in this subdomain.</li>\n<li><em>Bayesian/GP error gates calling the oracle only on novel environments during MD.</em> VASP MLFF (&gt;99% skip, many follow-up systems: zirconia, perovskites) and FLARE (&lt;100–250 calls) agree quantitatively.</li>\n<li><em>GP-surrogate gating of path/saddle searches.</em> Jónsson group&#39;s 2017–2020 series (order-of-magnitude fewer evaluations) was independently re-implemented with the same 3–10x conclusion by GPR_calculator (2025).</li>\n<li><em>Δ-ML economics.</em> Ramakrishnan/von Lilienfeld 2015 (1–10% labels) is independently echoed at materials scale by the VASP group&#39;s RPA-ΔML (168 high-level structures, &lt;150k CPU h) and by ANI-1ccx at molecular scale.</li>\n</ul>\n<p><strong>Single-group or young claims (real numbers, not yet replicated):</strong></p>\n<ul>\n<li>UDD-AL uncertainty-biased sampling (LANL, 2023) — conceptually mirrored by HAL (Cambridge, 2023, citing it), so the <em>idea</em> has two-lineage support, but the glycine/acetylacetone numbers are single-study.</li>\n<li>Annevelink–Viswanathan&#39;s statistical cutoff (2023) — important because it <em>quantifies when ensemble σ fails</em>; single study, two datasets.</li>\n<li>UBMD calibrated gradient UQ (Zaverkin/Kästner, 2024) — single group, strong systems (MIL-53(Al)).</li>\n<li>Evidential gates (eIP 2024, e²IP 2026; Fudan/DP-Technology lineage) — real AL/UDD demos but gate-economics numbers are qualitative.</li>\n<li>Flexible CP calibration on MACE-MP-0 (2025) — preprint.</li>\n</ul>\n<p><strong>Marketing-adjacent (true but inflated framing):</strong></p>\n<ul>\n<li>Gubaev&#39;s &quot;3–4 orders of magnitude&quot; folds surrogate relaxation speed into the DFT-call story.</li>\n<li>AdsorbML&#39;s &quot;2000x&quot; is relative to a Random+Heuristic DFT baseline, not to tuned DFT practice; the honest, robust number is the 87% success / 13% paid-tail measurement.</li>\n<li>&quot;Negligible-overhead UQ&quot; claims (eIP, flexible CP) never price the gate in DFT calls saved — a gap the field leaves open.</li>\n</ul>\n<p><strong>Notably thin (honest assessment):</strong> true <em>production-MD</em> abstention rates (post-training, the fraction of deployed inference sent back to DFT) are almost never reported; nearly all published fractions are from training campaigns. AdsorbML&#39;s ~13% verify tail and the on-the-fly &quot;calls become infrequent after 3–10 ps&quot; observation (FLARE) are the only production-mode numbers I could verify. Conformal/evidential gates likewise have <strong>no published DFT-call budgets</strong> — they sell calibration, not priced abstention.</p>\n<h2 id=\"4-openings-for-lupine\">4. Openings for Lupine</h2><ol>\n<li><strong>Union anchors are a genuinely unpriced quantity.</strong> Every published budget above is per-model/per-potential (DP-GEN 0.0044%, FLARE 250 calls, zirconia 168 RPA points). No paper measures DFT-call amortization across <em>multiple architecturally distinct uMLIPs</em> sharing an oracle. Lupine&#39;s 154-vs-558 union-anchor result (72.4% fewer evaluations), as far as I can verify, is the first cross-model sharing economy number in this literature — publishable as the missing &quot;multi-tenant abstention&quot; baseline, directly comparable against the per-model numbers in §2.</li>\n<li><strong>Barrier-local selective DFT, priced against GP-NEB.</strong> The only oracle-call economics for saddle searches are GP-surrogate NEB (10x, Jónsson) and its 2025 replication (3–10x). Both retrain a global surrogate during the search; neither <em>gates</em> evaluations by physical-law theorems. A Lupine risk-coverage curve in units of &quot;barrier error (meV) vs DFT anchors per path&quot; — with the theorem gate deciding <em>which</em> anchor is provably needed — would be the first abstention pricing for NEB and slots next to Koistinen&#39;s number as a competitor baseline.</li>\n<li><strong>Fix the documented failure mode of ensemble gates with certified accept regions.</strong> Annevelink–Viswanathan (and the 2024 spatially-resolved-uncertainty ChemRxiv study, accessed this session) show raw ensemble σ decorrelates from true error; current gates compensate with hand-tuned thresholds (e.g., Kulichenko&#39;s ρ = 0.35 kcal/mol/√N, chosen because the literature default was &quot;too low&quot;). A Lean-formalized theorem gate that certifies the accept region (rather than thresholding a heuristic σ) is exactly the repair this literature is implicitly asking for, and UDD-AL&#39;s own text flags automatic threshold selection as an open problem.</li>\n<li><strong>Runtime correction converts training-time abstention into production-time abstention.</strong> The measured production gap (§3) exists because FLARE/VASP gates decide &quot;retrain&quot; — their abstention ends when training ends. Lupine&#39;s gate decides &quot;correct this prediction now,&quot; which is the missing production-mode measurement: an abstention rate for deployed uMLIP MD (fraction of steps/structures corrected by sparse DFT anchors at runtime), with Δ-ML-style budgets (1–10% labels, 168-point high-level sets) as the pricing precedent.</li>\n</ol>\n<p>[UNVERIFIED items flagged inline: Tran &amp; Ulissi&#39;s exact DFT-call factor (paywalled full text); Botu &amp; Ramprasad&#39;s speedup factor; Schran committee&#39;s labeling budget; Hodapp &amp; Shapeev &quot;in operando active learning&quot; (MLST 1, 045005, 2020 — citation verified, numbers not accessed). All other quantitative claims above were read this session at the access level stated.]</p>\n"}