{"id":"savings-active-learning","title":"Active Learning for Interatomic Potentials as a DFT-Call Economizer","subtitle":"Active-learning loops for training interatomic potentials, priced in DFT calls avoided.","category":"references","tags":["literature-review","savings-stack","active-learning","mlip"],"source":"articles/lit-review/savings-active-learning.md","lang":"en","words":2304,"readMinutes":10,"toc":[{"depth":2,"text":"1. Executive result","id":"1-executive-result"},{"depth":2,"text":"2. Savings table","id":"2-savings-table"},{"depth":2,"text":"3. Proven vs claimed","id":"3-proven-vs-claimed"},{"depth":2,"text":"4. Openings for Lupine","id":"4-openings-for-lupine"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> explore agent <code>agent-14</code> (director-commissioned deep research, swarm of 7, 2026-07-21) — materialized verbatim, then editorially corrected only to replace the retracted 624/132 union-anchor figures with the reproducible primary-record values (558/154). Quantitative literature claims are as reported by the research agent from sources it accessed; see citations inline. Citation-verification pass pending before any external publication.</p>\n</blockquote>\n<h1 id=\"chapter-digest-active-learning-for-interatomic-potentials-as-a-dft-call-economizer\">Chapter Digest: Active Learning for Interatomic Potentials as a DFT-Call Economizer</h1><p><strong>Scope note.</strong> This subdomain is <em>not</em> thin — it is one of the most densely evidenced compute-savings areas in atomistic simulation. Every number below was read this session from the cited arXiv abstract page or full text (marked accordingly); journal-only items were cross-checked via Crossref/ACS metadata. Evidence cut: 2026-07-21.</p>\n<h2 id=\"1-executive-result\">1. Executive result</h2><ul>\n<li><strong>Uncertainty/committee-triggered labeling reduces DFT single-points by 2–4 orders of magnitude versus dense sampling, and this is replicated across five independent ecosystems.</strong> VASP&#39;s on-the-fly Bayesian MLFF bypasses &quot;&gt;99% of the first-principles calculations&quot; during force-field generation (<a href=\"https://arxiv.org/abs/1904.12961\">Jinnouchi, Karsai, Kresse, PRB 100, 014105 (2019), arXiv:1904.12961v3, abstract</a>); DP-GEN labeled only 0.0044% of <del>650M explored configurations for the Al–Mg alloy potential (<a href=\"https://arxiv.org/abs/1810.11890\">Zhang et al., PRM 3, 023804 (2019), arXiv:1810.11890v2, full text</a>); FLARE trains usable force fields with &quot;</del>100 DFT calculations&quot; per system (<a href=\"https://arxiv.org/abs/1904.02042\">Vandermause et al., npj Comput. Mater. 6, 20 (2020), arXiv:1904.02042v3, full text</a>).</li>\n<li><strong>The total DFT bill for a general-purpose single-element potential is now ~10²–10⁴ single-points, depending on accuracy target and phase-space coverage.</strong> ~100 calls for one-phase FLARE models; 7,646 labels (0.03% of 25M explored) for the general Cu DP-GEN potential covering 50 K–2T_m and 1–50,000 bar (<a href=\"https://arxiv.org/abs/1910.12690\">Zhang et al., CPC 253, 107206 (2020), arXiv:1910.12690v1, full text</a>); 814 reference calculations for a committee water model spanning liquid, ice phases, and the air–water interface <em>including</em> nuclear quantum effects (<a href=\"https://arxiv.org/abs/2006.01541\">Schran et al., JCP 153, 104105 (2020), arXiv:2006.01541v2, abstract</a>).</li>\n<li><strong>Query-by-committee delivers the cleanest measured data-reduction factor at fixed accuracy:</strong> ANI active learning matched ANI-1 on the COMP6 benchmark with <strong>10% of the data</strong> and &quot;vastly outperformed&quot; it with 25% — an order-of-magnitude economization against naive/random sampling (<a href=\"https://arxiv.org/abs/1801.09319\">Smith et al., JCP 148, 241733 (2018), arXiv:1801.09319v2, abstract</a>).</li>\n<li><strong>Rare events are the documented weak point of all mainstream triggers.</strong> DP-GEN&#39;s own authors state the committee indicator is &quot;only a sufficient, not a necessary, condition for poor performance&quot; (ensembles can agree and be wrong) (full text, arXiv:1810.11890); FLARE&#39;s GP epistemic uncertainty saturates and <em>underestimates</em> true error under large extrapolation (δ&gt;20% distortion) (full text, arXiv:1904.02042); standard MD-driven AL &quot;is prone to miss either rare events or extrapolative regions&quot; (<a href=\"https://arxiv.org/abs/2312.01416\">Zaverkin et al., npj Comput. Mater. 10, 83 (2024), arXiv:2312.01416v2, abstract</a>).</li>\n<li><strong>The savings frontier in 2025–2026 has shifted to (a) shrinking the cell, not just the count — small-cell AL gives up to 100× cost reduction (<a href=\"https://arxiv.org/abs/2504.07293\">Meng et al., arXiv:2504.07293v1, abstract</a>) — and (b) on-the-fly fine-tuning of foundation models (<a href=\"https://pubs.rsc.org/en/content/articlelanding/2026/dd/d5dd00392j\">Rensmeyer et al., Digital Discovery (2026), DOI 10.1039/D5DD00392J, abstract</a>)</strong> — i.e., the field is converging on exactly Lupine&#39;s design space.</li>\n</ul>\n<h2 id=\"2-savings-table\">2. Savings table</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Technique</th>\n<th>What cost it removes</th>\n<th>Measured savings factor (system/size context)</th>\n<th>Accuracy cost / failure mode</th>\n<th>Citation (what I read)</th>\n<th>Evidence strength</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Technique\">VASP MLFF on-the-fly Bayesian AL (2-/3-body + SOAP-like descriptors, Bayesian error on forces)</td>\n<td data-label=\"What cost it removes\">DFT calls during training MD; AIMD for production runs</td>\n<td data-label=\"Measured savings factor (system/size context)\">&gt;99% of FP calculations bypassed during FF generation; ~1000× simulation speedup; melting points of Al, Si, Ge, Sn, MgO</td>\n<td data-label=\"Accuracy cost / failure mode\">Only samples configurations the trajectory visits; validation vs thermodynamic perturbation needed</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1904.12961\">arXiv:1904.12961v3, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">VASP MLFF perspective (scaling of the above)</td>\n<td data-label=\"What cost it removes\">Same, generalized to large-scale MD</td>\n<td data-label=\"Measured savings factor (system/size context)\">&quot;most FP calculations bypassed… accelerated by several orders of magnitude&quot;</td>\n<td data-label=\"Accuracy cost / failure mode\">Qualitative; no per-system accounting</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://pubs.acs.org/doi/10.1021/acs.jpclett.0c01061\">J. Phys. Chem. Lett. 11, 6946 (2020), DOI 10.1021/acs.jpclett.0c01061, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Learn-on-the-fly (Li–Kermode–De Vita), GP extrapolation grade</td>\n<td data-label=\"What cost it removes\">Periodic re-evaluation of QM during MD</td>\n<td data-label=\"Measured savings factor (system/size context)\">Factor ~30 fewer QM calculations (as reported in the Podryabinkin–Shapeev full text)</td>\n<td data-label=\"Accuracy cost / failure mode\">Fixed-interval fallback needed; decision not geometry-driven</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://doi.org/10.1103/PhysRevLett.114.096405\">PRL 114, 096405 (2015)</a>, quoted in <a href=\"https://ar5iv.labs.arxiv.org/html/1611.09346\">arXiv:1611.09346v3 full text</a></td>\n<td data-label=\"Evidence strength\">moderate (number verified as quoted, original not read)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">MTP D-optimality active learning (MaxVol, extrapolation grade γ)</td>\n<td data-label=\"What cost it removes\">QM calls in MD/relaxation; guarantees interpolation-only prediction</td>\n<td data-label=\"Measured savings factor (system/size context)\">On-the-fly training with QM calls &quot;typically in the initial stage&quot;; no extrapolation attempted; threshold γ_th tunable</td>\n<td data-label=\"Accuracy cost / failure mode\">Restricted to linear-in-parameter models; γ is geometric, not an error bar</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1611.09346\">Podryabinkin &amp; Shapeev, CMS 140, 171 (2017), arXiv:1611.09346v3, abstract+full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ANI query-by-committee (ensemble disagreement)</td>\n<td data-label=\"What cost it removes\">Data needed for transferable organic-molecule NNP</td>\n<td data-label=\"Measured savings factor (system/size context)\">Parity with ANI-1 on COMP6 at 10% of data; dominant at 25% (~4–10× vs naive sampling)</td>\n<td data-label=\"Accuracy cost / failure mode\">QBC blind where committee shares bias; per-molecule DFT cost still ~10⁵–10⁶ points for universal sets</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1801.09319\">arXiv:1801.09319v2, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">DP-GEN concurrent learning (model deviation of 4-model DP ensemble)</td>\n<td data-label=\"What cost it removes\">Labels for broad (T,p) coverage of alloys</td>\n<td data-label=\"Measured savings factor (system/size context)\">Al: 6,489 / Mg: 3,689 / Al–Mg: 18,654 labels from ~650M explored (0.0044%); uniform accuracy 50–2000 K incl. defects, surfaces, liquid</td>\n<td data-label=\"Accuracy cost / failure mode\">Configs with deviation &gt; σ_hi (0.15 eV/Å) discarded as &quot;unphysical&quot; — rare high-error states can be dropped unlabeled; no rigorous indicator theory (authors&#39; own caveat)</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1810.11890\">arXiv:1810.11890v2, full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">DP-GEN platform (scheduler/dispatcher)</td>\n<td data-label=\"What cost it removes\">Human labor + labels for general-purpose potentials</td>\n<td data-label=\"Measured savings factor (system/size context)\">Cu: 7,646 labels (0.03%) from 25M explored, 48 iterations, 0–2T_m, 1 bar–50 kbar</td>\n<td data-label=\"Accuracy cost / failure mode\">Same trust-band caveat (σ_lo 0.05, σ_hi 0.20 eV/Å); labeling caps at 300 configs/iteration</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1910.12690\">arXiv:1910.12690v1, full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Committee NNP (c-NNP, committee disagreement monitored/biased in MD)</td>\n<td data-label=\"What cost it removes\">Ab initio set for water incl. NQE</td>\n<td data-label=\"Measured savings factor (system/size context)\">814 reference calculations total for liquid water (ambient + elevated T,p), ice phases, air–water interface, with path-integral NQE</td>\n<td data-label=\"Accuracy cost / failure mode\">Committee = multiple models → training cost multiplied; state-point-specific, not universal</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2006.01541\">arXiv:2006.01541v2, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Committee AL for aqueous interfaces (CP2K/NNP)</td>\n<td data-label=\"What cost it removes\">Model construction effort for complex aqueous systems</td>\n<td data-label=\"Measured savings factor (system/size context)\">Automated AL after one initial AIMD run; validated structure/dynamics for ions, TiO2/water, confined water</td>\n<td data-label=\"Accuracy cost / failure mode\">Committee-size and validation protocol required; single-state-point models</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2106.00048\">Schran et al., PNAS 118, e2110077118 (2021), arXiv:2106.00048v2, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">FLARE (low-dim GP, Bayesian epistemic/noise uncertainty)</td>\n<td data-label=\"What cost it removes\">Training data for on-the-fly force fields; AIMD cost for rare-event MD</td>\n<td data-label=\"Measured savings factor (system/size context)\">~100 DFT calls/system (C:107, Si:133, Al₂O₃:16, NiTi:18, BN:237, AgI:39); Al vacancy 1 ns training run &gt;300× faster than equivalent AIMD; HEA: GP matched NNs trained on &gt;10× more structures</td>\n<td data-label=\"Accuracy cost / failure mode\">GP epistemic uncertainty has an upper bound — true error underestimated for δ&gt;20% distortions; threshold tied to noise σ_n needs care at phase changes</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/1904.02042\">arXiv:1904.02042v3, full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">FLARE++ (sparse GP + ACE descriptors, mapped to polynomial)</td>\n<td data-label=\"What cost it removes\">Training data + inference cost for reactive force fields</td>\n<td data-label=\"Measured savings factor (system/size context)\">216 DFT calls (reactive Pt/H) + 24/4/6 (H₂, Pt surface, bulk) ≈ 250 total → reactive FF in 3 days; 2× faster than Pt/H ReaxFF; E_a 0.25(2) eV vs exp 0.23 eV</td>\n<td data-label=\"Accuracy cost / failure mode\">Single-system demonstration; extrapolation confidence grows outside training set but flagged only by GP variance</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://ar5iv.labs.arxiv.org/html/2106.01949\">Vandermause et al., Nat. Commun. 13, 5183 (2022), arXiv:2106.01949v1, full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GAP iterative learning (a-C)</td>\n<td data-label=\"What cost it removes\">Full AIMD melt–quench trajectories (~9,500 steps) per uncorrelated amorphous sample</td>\n<td data-label=\"Measured savings factor (system/size context)\">Protocol: preliminary GAP drives melt–quench; only 1+ DFT single-point per GAP trajectory retained; iteratively extended across densities 1.5–3.5 g/cm³</td>\n<td data-label=\"Accuracy cost / failure mode\">Early many-body-only model produced sub-Å atom aggregation (unphysical extrapolation); total DFT count not stated in main text [UNVERIFIED — given only in SI]</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://ar5iv.labs.arxiv.org/html/1611.03277\">Deringer &amp; Csányi, PRB 95, 094203 (2017), arXiv:1611.03277v1, full text</a></td>\n<td data-label=\"Evidence strength\">strong (protocol), [UNVERIFIED] total count</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ACE D-optimality vs ensemble (MaxVol extrapolation grade)</td>\n<td data-label=\"What cost it removes\">Ensemble training cost for ACE; rare-event discovery</td>\n<td data-label=\"Measured savings factor (system/size context)\">D-optimality indicator &quot;more computationally efficient&quot; than ensembles at comparable prediction; enables automated rare-event exploration incl. per-environment selection in large MD</td>\n<td data-label=\"Accuracy cost / failure mode\">Two query strategies compared, no single hard DFT-call factor reported</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2212.08716\">Lysogorskiy et al., PRM 7, 043801 (2023), arXiv:2212.08716v1, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">MLIP-3 per-atomic-environment AL (MTP)</td>\n<td data-label=\"What cost it removes\">Whole-configuration labeling in large simulations</td>\n<td data-label=\"Measured savings factor (system/size context)\">Active learning on atomic neighborhoods, enabling AL in large-scale runs where whole-frame selection is too coarse</td>\n<td data-label=\"Accuracy cost / failure mode\">Software/methods paper; efficiency claims qualitative</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2304.13144\">Podryabinkin et al., JCP 159, 114104 (2023), arXiv:2304.13144v3, abstract</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">FALCON OTF calculator (AGOX GPR, ASE)</td>\n<td data-label=\"What cost it removes\">DFT calls in long MD; retraining overhead via k-means clustered sub-models</td>\n<td data-label=\"Measured savings factor (system/size context)\">Al₃₂ melting, 250 ps: 500 DFT calls in first 25 ps, then only 69 in remaining 225 ps; H₂O-in-CNT (115 atoms): 151 DFT calls for 50,000 steps, &gt;300× vs AIMD (500 days → 38.4 h)</td>\n<td data-label=\"Accuracy cost / failure mode\">DFT-call count swings ~4× with threshold choice (0.10→0.20 eV: &gt;300→86 trainings); GPR retraining can exceed AIMD cost for large/hot/disordered systems (authors&#39; own warning)</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://www.nature.com/articles/s41524-025-01897-8\">Felis &amp; Dononelli, npj Comput. Mater. 12 (2026), DOI 10.1038/s41524-025-01897-8, full text</a></td>\n<td data-label=\"Evidence strength\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Small-cell AL (1–8 atom training cells)</td>\n<td data-label=\"What cost it removes\">Large-cell DFT cost during AL</td>\n<td data-label=\"Measured savings factor (system/size context)\">Up to 100× cost savings vs 54-atom-cell training; some runs &lt;120 core-hours (K, Na–K; solid–liquid interfaces, critical exponents reproduced)</td>\n<td data-label=\"Accuracy cost / failure mode\">Single group; two alkali-metal systems; generality unproven</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2504.07293\">Meng et al., Comput. Mater. Sci. 256, 113919 (2025), arXiv:2504.07293v1, abstract</a></td>\n<td data-label=\"Evidence strength\">moderate→strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Uncertainty-biased MD (energy-uncertainty bias forces + bias stress)</td>\n<td data-label=\"What cost it removes\">Missed rare events / extrapolative regions of plain AL-MD</td>\n<td data-label=\"Measured savings factor (system/size context)\">MLIPs &quot;similar or better accuracy than ensemble-based methods at lower computational cost&quot; (alanine dipeptide, MIL-53(Al))</td>\n<td data-label=\"Accuracy cost / failure mode\">Requires calibrated gradient uncertainties; new method, limited replication</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2312.01416\">arXiv:2312.01416v2, abstract</a></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">On-the-fly fine-tuning of foundation models (Bayesian NN UQ)</td>\n<td data-label=\"What cost it removes\">Training-data burden when specializing pretrained uMLIPs</td>\n<td data-label=\"Measured savings factor (system/size context)\">Automates fine-tuning at pre-specified accuracy; &quot;detects rare events such as transition states and samples them at an increased rate&quot;</td>\n<td data-label=\"Accuracy cost / failure mode\">No hard DFT-call factor in abstract; foundation models lack native UQ</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://pubs.rsc.org/en/content/articlelanding/2026/dd/d5dd00392j\">Rensmeyer et al., Digital Discovery (2026), DOI 10.1039/D5DD00392J, abstract</a></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">LOCAL (GNN + locality AL, catalyst screening)</td>\n<td data-label=\"What cost it removes\">DFT labels for stability screening of 611,648 DAC/NG structures</td>\n<td data-label=\"Measured savings factor (system/size context)\">16,704 labels (2.7%) → 0.15 eV test MAE</td>\n<td data-label=\"Accuracy cost / failure mode\">Screening task (not MD potential); single paper</td>\n<td data-label=\"Citation (what I read)\"><a href=\"https://arxiv.org/abs/2503.19445\">Yin et al., arXiv:2503.19445v5, abstract</a></td>\n<td data-label=\"Evidence strength\">moderate</td>\n</tr>\n</tbody></table></div><h2 id=\"3-proven-vs-claimed\">3. Proven vs claimed</h2><p><strong>Replicated and trusted (multi-group, multi-codebase):</strong></p>\n<ul>\n<li>Uncertainty/committee-triggered labeling cuts DFT single-points by <del>10×–1000× versus dense or random sampling at matched accuracy. Independently demonstrated by the ANI line (10% of data at parity, LANL/UNC), DP-GEN line (0.0044–0.03% label rates, Princeton/IAPCM/DeepModeling), VASP-MLFF line (&gt;99% bypass, Vienna/Toyota), FLARE line (</del>100 calls/system, Harvard/Bosch), and MTP line (Skoltech). These are production workflows with thousands of citations and open-source implementations (DeePMD-kit/DP-GEN, VASP≥6.3, FLARE, MLIP-3, ASE calculators).</li>\n<li>GP/Bayesian internal uncertainties track true error well <em>near</em> the training set (FLARE Fig. 2a–c; VASP Bayesian error) — sufficient as <em>triggers</em>, not as <em>guarantees</em>.</li>\n<li>Iterative database extension (GAP line, Cambridge) is the trusted manual ancestor of automated AL and documented its own early failure (sub-Å clumping) before hierarchical descriptors fixed it.</li>\n</ul>\n<p><strong>Single-paper or thin-replication claims:</strong></p>\n<ul>\n<li>FLARE&#39;s &quot;~100 DFT calls&quot; per system: now partially corroborated by FALCON (151 calls for H₂O/CNT; 569 for Al melting over 250 ps), so graduating toward proven — but both are GPR-family, and FALCON itself warns retraining can eat the savings for large/hot systems.</li>\n<li>Small-cell AL 100× (one 2025 paper, alkali metals only).</li>\n<li>Uncertainty-biased MD and Bayesian fine-tuning of foundation models (2024/2026, abstracts only here; no independent replication yet).</li>\n<li>Per-environment AL (MLIP-3) and ACE MaxVol-vs-ensemble equivalence: methods papers without hard per-system DFT-call factors.</li>\n</ul>\n<p><strong>Marketing / handle-with-care:</strong></p>\n<ul>\n<li>&quot;Orders of magnitude&quot; in perspective pieces (JPCL 2020 included) never define the baseline; the honest unit is <em>DFT single-points to reach a stated force RMSE on a stated validation set</em>, and almost no two papers use the same one.</li>\n<li>Reported label fractions (0.0044%) ignore exploration cost — cheap per step but run over 10⁷–10⁸ configurations, and ignore the 4× ensemble training overhead.</li>\n<li>Threshold sensitivity is real: FALCON&#39;s own numbers show a 4× swing in DFT calls between accuracy thresholds of 0.10 and 0.20 eV. Any &quot;DFT calls needed&quot; figure without its threshold is not comparable across papers.</li>\n<li>The a-C total training-set size is widely quoted second-hand but is not in the main text [UNVERIFIED — stated only in the paper&#39;s SI, not accessed].</li>\n</ul>\n<h2 id=\"4-openings-for-lupine\">4. Openings for Lupine</h2><ol>\n<li><strong>Theorem-gating fills the documented &quot;confident-but-wrong&quot; blind spot of every learned-UQ trigger.</strong> DP-GEN&#39;s authors concede the committee indicator is sufficient-but-not-necessary (ensemble can agree and be wrong, arXiv:1810.11890 full text); FLARE&#39;s GP variance saturates and <em>underestimates</em> error under strong extrapolation (arXiv:1904.02042 full text); Tan et al. benchmarked that no UQ scheme consistently flags OOD (<a href=\"https://arxiv.org/abs/2305.01754\">npj Comput. Mater. 2023, arXiv:2305.01754v1, abstract</a>). A Lean-formalized physical-law check (symmetry, conservation, curvature-sign at saddles) is an independent, non-learned oracle that can veto predictions exactly where committees and GP variances are structurally blind — Lupine should frame theorem-gating as the <em>necessary-condition</em> complement to AL&#39;s sufficient-condition triggers.</li>\n<li><strong>Union anchors are an orthogonal savings axis no AL framework addresses.</strong> AL economizes labels <em>within</em> one model&#39;s training; nothing in this literature deduplicates DFT evaluations <em>across</em> models. Lupine&#39;s measured 154 shared vs 558 naive evaluations (72.4% fewer on a 29-path panel) stacks multiplicatively with AL&#39;s per-model savings: run FLARE/DP-GEN-style uncertainty triggers, but only label the union of cross-model disagreement. The same logic extends small-cell AL (Meng et al. 2025, 100×): Lupine&#39;s sparse anchors can be evaluated as small cells around the predicted saddle for another ~2 orders of magnitude, turning the AL &quot;label count&quot; game into a &quot;label count × cell volume&quot; game.</li>\n<li><strong>Barrier-focused sparse anchoring vs rare-event MD learning.</strong> FLARE&#39;s flagship rare-event demonstration (Al vacancy) needed a 1 ns on-the-fly run to learn one migration mechanism; Lupine&#39;s 4–6 DFT images near the model-predicted saddle gets the barrier directly. The 2026 foundation-model fine-tuning line (Rensmeyer) independently identifies transition states as the high-value rare events worth oversampling — Lupine can position sparse/union anchoring + theorem-verified saddle geometry as the deterministic, retraining-free answer to the same need.</li>\n<li><strong>Replace &quot;label-and-retrain&quot; with &quot;verify-and-patch&quot; to kill AL&#39;s documented cost spike.</strong> FALCON&#39;s honest accounting shows GPR retraining can dominate and even exceed AIMD cost at tight thresholds or high T (full text). Lupine&#39;s runtime correction patches uMLIP errors against physical theorems without retraining the base model — directly attacking the cost component (retraining overhead) that AL papers systematically under-report, while keeping a DFT ledger only for theorem-violating regions.</li>\n</ol>\n<p><strong>Caveats for the chapter editors.</strong> (i) Cross-paper DFT-call comparisons are apples-to-oranges — thresholds, system sizes, and phase-space coverage differ; I flagged context per row. (ii) Two numbers I could not verify are marked [UNVERIFIED] (a-C total database size; Li et al. 2015&#39;s factor-30 is verified only as quoted in the Podryabinkin–Shapeev full text). (iii) The MAPbI₃ PRL (arXiv:1903.09613) and stanene npj paper (arXiv:2008.11796) were verified at abstract level only — cited for existence/claims, not for per-system DFT counts. (iv) The Zaverkin (arXiv:2312.01416) &quot;existing methods miss rare events&quot; statement is the field&#39;s own most quotable failure admission — useful rhetorically, but it is an abstract-level claim of a 2024 single paper.</p>\n"}