{"id":"savings-delta-multifidelity","title":"Δ-Learning, Multi-Fidelity, and Transfer-Learning Corrections","subtitle":"Correction layers that learn the gap between cheap and expensive electronic-structure methods.","category":"references","tags":["literature-review","savings-stack","delta-learning","multi-fidelity","transfer-learning"],"source":"articles/lit-review/savings-delta-multifidelity.md","lang":"en","words":2420,"readMinutes":11,"toc":[{"depth":2,"text":"1. Executive result","id":"1-executive-result"},{"depth":2,"text":"2. Savings table","id":"2-savings-table"},{"depth":2,"text":"3. Proven vs claimed","id":"3-proven-vs-claimed"},{"depth":2,"text":"4. Openings for Lupine","id":"4-openings-for-lupine"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> explore agent <code>agent-15</code> (director-commissioned deep research, swarm of 7, 2026-07-21) — materialized verbatim. Quantitative claims are as reported by the research agent from sources it accessed; see citations inline. Citation-verification pass pending before any external publication.</p>\n</blockquote>\n<h1 id=\"chapter-section-learning-multi-fidelity-and-transfer-learning-corrections-insider-39-s-digest\">Chapter Section: Δ-Learning, Multi-Fidelity, and Transfer-Learning Corrections — Insider&#39;s Digest</h1><p><em>Evidence cut: 2026-07-21. All citations below were accessed this session; each is annotated with what was actually read (abstract / full text / snippet).</em></p>\n<h2 id=\"1-executive-result\">1. Executive result</h2><ul>\n<li><strong>Fine-tuning a pretrained universal MLIP is now the dominant compute-savings move in this subdomain, and the measured data-efficiency numbers are large and independently replicated across groups.</strong> The single most striking verified result: Microsoft&#39;s MatterSim fine-tuned to revPBE0-D3 accuracy for liquid water with <strong>30 high-fidelity configurations</strong> matching a from-scratch model trained on <strong>900</strong> (30× fewer labels; abstract headline &quot;up to 97% reduction in data requirements&quot;) — <a href=\"https://arxiv.org/abs/2405.04967\">MatterSim, arXiv:2405.04967v2</a> (full text read). DPA-2 reports <strong>1–2 orders of magnitude</strong> less downstream data across 18 pretraining datasets (2 orders on the H₂O-PBE0TS task specifically) — <a href=\"https://arxiv.org/abs/2312.15492\">DPA-2, arXiv:2312.15492v2</a> (full text read).</li>\n<li><strong>The Δ-learning stack (cheap baseline + learned residual) reliably buys 10–100× on high-fidelity labels when the correction is smooth.</strong> Multi-fidelity CQML reaches <del>1 kcal/mol chemical accuracy on atomization energies with **</del>100 CCSD(T) labels instead of thousands**, shifting the data burden to cheap fidelities — <a href=\"https://arxiv.org/abs/1808.02799\">Zaspel et al., arXiv:1808.02799v2 / JCTC 15:1546 (2019), DOI 10.1021/acs.jctc.8b00832</a> (abstract read). The original Δ-ML result: models trained on <strong>1–10% of 134k molecules</strong> reproduce enthalpies of the remainder at DFT accuracy — <a href=\"https://arxiv.org/abs/1503.04987\">Ramakrishnan et al., arXiv:1503.04987 / JCTC 11:2087 (2015), DOI 10.1021/acs.jctc.5b00099</a> (abstract read).</li>\n<li><strong>Transfer learning DFT→gold-standard works at ecosystem scale, not just toy scale.</strong> ANI-1ccx: 5.2M DFT points → retrain on <del>500k DLPNO-CCSD(T)/CBS points; retraining takes **</del>30 min vs ~4 h** for from-scratch training per network, and the transfer-learned model beats a model trained on the same coupled-cluster data alone by <strong>23% RMSD</strong> — <a href=\"https://www.nature.com/articles/s41467-019-10827-4\">Smith et al., Nat. Commun. 10:2903 (2019), DOI 10.1038/s41467-019-10827-4</a> (full text read). AIQM1 stacks Δ-learning (SQM→DFT, 4.6M points) + transfer learning (→CCSD(T)*/CBS, 0.5M points) to reach near-G4 accuracy (heats of formation MAD 0.9 kcal/mol vs 2.6 for the semiempirical parent) at semiempirical cost — <a href=\"https://www.nature.com/articles/s41467-021-27340-2\">Zheng et al., Nat. Commun. 12:7022 (2021), DOI 10.1038/s41467-021-27340-2</a> (full text read).</li>\n<li><strong>For Lupine&#39;s home territory (barriers/NEB), the closest published competitor result:</strong> fine-tuned CHGNet cuts Li-ion migration-barrier MAE from ~0.23–0.24 eV (pretrained) to <strong>0.07–0.09 eV</strong>, with claimed 100–1000× speedup over DFT NEB, using transition-state-focused fine-tuning sets — <a href=\"https://arxiv.org/abs/2507.02334\">Lian et al., arXiv:2507.02334 / J. Mater. Chem. A 13:34918 (2025), DOI 10.1039/D5TA05355B</a> (abstract read; specific MAE values corroborated in the RSC full text and secondary summaries).</li>\n<li><strong>Fine-tuning has documented failure modes that runtime correction avoids</strong>: fine-tuning on relaxation-heavy data <em>reintroduces systematic softening</em> of the PES — <a href=\"https://arxiv.org/abs/2410.12771\">OMat24, arXiv:2410.12771v2</a> (full text read); &quot;naive&quot; fine-tuning on single MD trajectories fails out-of-distribution — <a href=\"https://arxiv.org/abs/2603.10159\">Wong et al., arXiv:2603.10159v2 / JCTC, DOI 10.1021/acs.jctc.6c00425</a> (abstract read); and pretrained representations carry chemistry-specific biases (metal–sulfur underrepresentation) that Δ-models on the model&#39;s own embedding must patch — <a href=\"https://arxiv.org/abs/2502.21179\">Hammer, arXiv:2502.21179v2</a> (abstract read).</li>\n</ul>\n<h2 id=\"2-savings-table\">2. Savings table</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Technique</th>\n<th>What cost it removes</th>\n<th>Measured savings (system/size context)</th>\n<th>Accuracy cost / failure mode</th>\n<th>Citation (accessed as)</th>\n<th>Evidence</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Technique\">Δ-ML kernel correction (semiempirical→DFT)</td>\n<td data-label=\"What cost it removes\">High-level QC labels for training</td>\n<td data-label=\"Measured savings (system/size context)\">Chemical accuracy on enthalpies of 134k organic molecules training on 1–10% of them (GDB-9-derived set)</td>\n<td data-label=\"Accuracy cost / failure mode\">Residual must be smooth/small; kernel scaling O(N³); energies only in original</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/1503.04987\">Ramakrishnan 2015, arXiv:1503.04987; JCTC 11:2087, DOI 10.1021/acs.jctc.5b00099</a> (abstract)</td>\n<td data-label=\"Evidence\">strong (foundational, widely replicated)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Multi-fidelity CQML (HF/MP2/CCSD(T) stack)</td>\n<td data-label=\"What cost it removes\">~90–99% of gold-standard labels</td>\n<td data-label=\"Measured savings (system/size context)\">~100 CCSD(T)/cc-pVDZ labels instead of thousands for ~1 kcal/mol atomization-energy accuracy on ~7k molecules</td>\n<td data-label=\"Accuracy cost / failure mode\">Needs nested fidelity structure; KRR cost grows with multi-level data</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/1808.02799\">Zaspel 2019, arXiv:1808.02799v2; JCTC 15:1546, DOI 10.1021/acs.jctc.8b00832</a> (abstract)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Transfer learning of NNP (DFT→CCSD(T)*/CBS)</td>\n<td data-label=\"What cost it removes\">High-fidelity data + training time</td>\n<td data-label=\"Measured savings (system/size context)\">ANI-1ccx: 500k CC labels on top of 5.2M DFT; 23% lower RMSD than training on the CC data alone; retrain 30 min vs 4 h per net; torsion MAD halved vs DFT-trained model (0.23 vs 0.47 kcal/mol median)</td>\n<td data-label=\"Accuracy cost / failure mode\">Still needs ~10⁵ high-level points for broad chemistry; paper itself notes Δ-learning reaches similar accuracy but doubles inference cost</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://www.nature.com/articles/s41467-019-10827-4\">Smith 2019, Nat. Commun. 10:2903, DOI 10.1038/s41467-019-10827-4</a> (full text)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Δ + transfer-learning hybrid on semiempirical baseline</td>\n<td data-label=\"What cost it removes\">DFT/CC cost at inference</td>\n<td data-label=\"Measured savings (system/size context)\">AIQM1: near-G4/G4MP2 heats of formation (MAD 0.9 kcal/mol, CHNO set; parent SQM 2.6) at semiempirical speed; ~1000×+ vs DFT geometry optimization (14 s vs 31 min×32 cores for C₆₀)</td>\n<td data-label=\"Accuracy cost / failure mode\">Neutral closed-shell H/C/N/O only; fails where training data sparse (H₂ error −2.9 kcal/mol)</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://www.nature.com/articles/s41467-021-27340-2\">Zheng 2021, Nat. Commun. 12:7022, DOI 10.1038/s41467-021-27340-2</a> (full text)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Δ-ML PES correction (DFT→CCSD(T), PIP fits)</td>\n<td data-label=\"What cost it removes\">CC gradients/labels for full PES</td>\n<td data-label=\"Measured savings (system/size context)\">DFT-based PES elevated to near-CCSD(T) for molecules up to 15 atoms (H₃O⁺…acetylacetone, tropolone; ethanol study across PBE/M06/M06-2X/PBE0+MBD); improvement even in <em>gradients</em> without CC gradients in fit</td>\n<td data-label=\"Accuracy cost / failure mode\">Single-molecule PESs; needs dense low-level sampling first</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://doi.org/10.1063/5.0038301\">Nandi 2021, JCP 154:051102, DOI 10.1063/5.0038301</a> (citation+snippet); <a href=\"https://arxiv.org/abs/2407.20050\">Nandi 2024, arXiv:2407.20050; JCTC 20:8807, DOI 10.1021/acs.jctc.4c00977</a> (abstract)</td>\n<td data-label=\"Evidence\">strong (Bowman-group series, replicated)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Multi-fidelity graph networks (PBE→expt)</td>\n<td data-label=\"What cost it removes\">Scarce experimental/high-fidelity labels</td>\n<td data-label=\"Measured savings (system/size context)\">22–45% lower MAE on experimental band gaps of ordered+disordered materials using PBE gaps as low fidelity</td>\n<td data-label=\"Accuracy cost / failure mode\">Property-prediction, not PES; gains shrink as high-fidelity data grows</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2005.04338\">Chen et al. 2021, arXiv:2005.04338v3; Nature Computational Science 1:46–53, DOI 10.1038/s43588-020-00002-x</a> (abstract)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Implicit delta multi-task NNP (semiempirical→DFT)</td>\n<td data-label=\"What cost it removes\">High-fidelity labels</td>\n<td data-label=\"Measured savings (system/size context)\">IDLe matches single-fidelity baseline with up to <strong>50× less</strong> high-fidelity data (SPICE-like molecular datasets; 11M semiempirical points released)</td>\n<td data-label=\"Accuracy cost / failure mode\">Single preprint; multi-head architecture required; Δ implicit, no physics handle at inference</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2412.06064\">Tossou et al. 2024, arXiv:2412.06064v1</a> (abstract)</td>\n<td data-label=\"Evidence\">claim (single preprint)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">uMLIP fine-tuning to higher level of theory</td>\n<td data-label=\"What cost it removes\">From-scratch training data &amp; time</td>\n<td data-label=\"Measured savings (system/size context)\">MatterSim water: 30 revPBE0-D3 configs ≈ from-scratch-900 (97% fewer labels; reproduces RDF/ADF + diffusion within 20% of expt); Li₂B₁₂H₁₂ active learning: 15% of data</td>\n<td data-label=\"Accuracy cost / failure mode\">Numbers from model authors; best-case system; fine-tune inherits pretrained bias</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2405.04967\">MatterSim, arXiv:2405.04967v2</a> (full text)</td>\n<td data-label=\"Evidence\">strong (but single-team)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Multi-task pretrain → fine-tune (DPA-2)</td>\n<td data-label=\"What cost it removes\">Downstream DFT labels</td>\n<td data-label=\"Measured savings (system/size context)\">1–2 orders of magnitude less downstream data across tasks (alloys, SSE, ferroelectrics, drug molecules); 2 orders on H₂O-PBE0TS; distillation recovers ~2 orders inference speed</td>\n<td data-label=\"Accuracy cost / failure mode\">Gains shrink as downstream data grows; needs overlap between pretraining and target chemistry</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2312.15492\">DPA-2, arXiv:2312.15492v2</a> (full text)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Cross-dataset fine-tuning of big pretrained models (OMat24→MPtrj)</td>\n<td data-label=\"What cost it removes\">Re-training on 110M structures per user</td>\n<td data-label=\"Measured savings (system/size context)\">OMat24-pretrained eSEN/eqV2 fine-tuned to MPtrj/sAlex: Matbench Discovery F1 0.925, energy-above-hull MAE 18 meV/atom (best at time of writing)</td>\n<td data-label=\"Accuracy cost / failure mode\"><strong>Fine-tuning on relaxation data reintroduces systematic softening</strong> (own analysis)</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2410.12771\">OMat24, arXiv:2410.12771v2</a> (full text)</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">TS-focused uMLIP fine-tuning for barriers (CHGNet)</td>\n<td data-label=\"What cost it removes\">DFT NEB/AIMD cost at scale</td>\n<td data-label=\"Measured savings (system/size context)\">Migration-barrier MAE 0.23–0.24 → 0.07–0.09 eV on Li-ion conductors; HT-NEB screening of quaternary Li compounds; claimed 100–1000× vs DFT NEB; &lt;1k DFT structures per model</td>\n<td data-label=\"Accuracy cost / failure mode\">Fine-tune set must contain transition states (chicken-and-egg: pretrained NEB generates them); 0.07–0.09 eV residual error still large for kinetics</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2507.02334\">Lian 2025, arXiv:2507.02334; JMCA 13:34918, DOI 10.1039/D5TA05355B</a> (abstract + RSC text snippet)</td>\n<td data-label=\"Evidence\">moderate-strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Small-set uMLIP fine-tuning for phonons/EXAFS (CHGNet)</td>\n<td data-label=\"What cost it removes\">Compound-specific DFT datasets</td>\n<td data-label=\"Measured savings (system/size context)\">~100 DFT structures fine-tune CHGNet to DFT-level phonons/thermal disorder for 2H-WS₂/MoS₂; even 1 structure partially viable</td>\n<td data-label=\"Accuracy cost / failure mode\">Compound-specific; force MAE floor ~100 meV/Å</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2509.08498\">Žguns 2025, arXiv:2509.08498v1</a> (abstract)</td>\n<td data-label=\"Evidence\">moderate-strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Fine-tune vs from-scratch, controlled comparisons (MACE)</td>\n<td data-label=\"What cost it removes\">Training data + convergence time</td>\n<td data-label=\"Measured savings (system/size context)\">Fine-tuned MACE beats from-scratch at every data fraction (10–100%) on LGPS (final force RMSE ~15 meV/Å); Mo elastic constants: C11 error 45.9%→2.6%, C44 56.9%→14.8%, ν 48.3%→3.5%; beats uncertainty-filtered from-scratch on Si OOD set</td>\n<td data-label=\"Accuracy cost / failure mode\">Single-GPU tutorial-scale; hyperparameter-sensitive; forgetting needs mitigations (num_samples_pt, multi-head replay)</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/html/2506.21935v2\">Liu/Zeng et al. 2025, arXiv:2506.21935v2</a> (full text)</td>\n<td data-label=\"Evidence\">moderate (preprint tutorial, very thorough)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Architecture-agnostic fine-tune sweep (MACE/GRACE/SevenNet/MatterSim/ORB)</td>\n<td data-label=\"What cost it removes\">System-specific ab initio accuracy</td>\n<td data-label=\"Measured savings (system/size context)\">Fine-tuning improves force accuracy 5–15×, energy accuracy 2–4 orders of magnitude, across 7 compounds, using equidistantly sampled short AIMD frames</td>\n<td data-label=\"Accuracy cost / failure mode\">Single preprint; &quot;near-ab initio&quot; claim system-dependent</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2511.05337\">Hänseroth 2025, arXiv:2511.05337v1</a> (abstract)</td>\n<td data-label=\"Evidence\">claim→moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPR Δ-model on uMLIP&#39;s own embeddings</td>\n<td data-label=\"What cost it removes\">Re-fitting / full fine-tuning</td>\n<td data-label=\"Measured savings (system/size context)\">CHGNet errors on sulfur-overlayer global optimization (Cu/Ag/Au(111)) corrected by sparse-GPR Δ-model trained on small DFT sets; same bias found in MACE-MP-0, SevenNet-0, ORB-v2</td>\n<td data-label=\"Accuracy cost / failure mode\">Correction is global/static, not query-specific; needs active-learning loop for GO</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2502.21179\">Hammer 2025, arXiv:2502.21179v2</a> (abstract); <a href=\"https://arxiv.org/pdf/2507.18485\">Pitfield 2026, arXiv:2507.18485; PCCP 28:912, DOI 10.1039/D5CP04302F</a> (abstract snippet)</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Periodic (iterative) fine-tuning</td>\n<td data-label=\"What cost it removes\">DFT labels while avoiding OOD failure</td>\n<td data-label=\"Measured savings (system/size context)\">Periodic fine-tuning generalizes where &quot;naive&quot; single-trajectory fine-tuning fails OOD (uMLIP MD bias study)</td>\n<td data-label=\"Accuracy cost / failure mode\">Fine-tuning can <em>inject</em> model bias; needs uncertainty monitoring (Q-residual)</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2603.10159\">Wong 2026, arXiv:2603.10159v2; JCTC, DOI 10.1021/acs.jctc.6c00425</a> (abstract)</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Fine-tune across DFT settings (UMA)</td>\n<td data-label=\"What cost it removes\">Re-generating data at consistent settings</td>\n<td data-label=\"Measured savings (system/size context)\">UMA models fine-tuned on MPtrj/sAlex to align DFT settings for materials evals; single un-fine-tuned model ≈ specialized models (abstract claim)</td>\n<td data-label=\"Accuracy cost / failure mode\">No public data-efficiency numbers for fine-tuning in accessed text</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2506.23971\">UMA, arXiv:2506.23971v1/v2</a> (abstract + PDF snippet)</td>\n<td data-label=\"Evidence\">moderate (fine-tune numbers) / strong (zero-shot claim)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Data-cost benchmarking of Δ vs MFML</td>\n<td data-label=\"What cost it removes\">Misallocated QC budget</td>\n<td data-label=\"Measured savings (system/size context)\">MFML beats Δ-ML when many predictions needed; Δ-ML (and MFΔML) wins for few predictions (QeMFi benchmark, 3 properties)</td>\n<td data-label=\"Accuracy cost / failure mode\">Kernel methods; small molecules</td>\n<td data-label=\"Citation (accessed as)\"><a href=\"https://arxiv.org/abs/2410.11391\">Vinod &amp; Zaspel 2024/25, arXiv:2410.11391v3</a> (abstract)</td>\n<td data-label=\"Evidence\">strong (guidance, not a technique)</td>\n</tr>\n</tbody></table></div><h2 id=\"3-proven-vs-claimed\">3. Proven vs claimed</h2><p><strong>Replicated and trusted (multi-group, multi-year):</strong></p>\n<ul>\n<li>Δ-learning of smooth level-of-theory residuals. From Ramakrishnan 2015 through the Bowman-group PES series (2021–2024, five functionals, up to 15-atom molecules, JCTC 2024 ethanol study) to AIQM1 at production scale — the 10–100× label savings are real when the residual is small and smooth. This is the most battle-tested idea in the subdomain.</li>\n<li>Transfer learning DFT→coupled-cluster for organic molecules. ANI-1ccx (2019) → AIQM1 (2021) → AIMNet/ANI ecosystem; the ~one-order-of-magnitude data saving and the &quot;retrain in minutes not hours&quot; training-cost saving are consistent across papers. Note the ANI-1ccx paper itself reports that plain Δ-learning matches transfer learning in accuracy — at 2× inference cost (full text, Methods).</li>\n<li>Fine-tuning pretrained uMLIPs to a specific chemistry. By mid-2026 this is independently demonstrated by at least six groups across five architecture families (MACE: tutorial arXiv:2506.21935; CHGNet: Lian, Žguns; MatterSim: Yang et al.; cross-architecture: Hänseroth arXiv:2511.05337; halide SSEs: arXiv:2510.09861; DPA line: DPA-2). Convergence on &quot;~10²–10³ structures per system&quot; as the practical fine-tune budget (100 for phonons per Žguns; 30–900 for level-of-theory uplift per MatterSim; &lt;1k for barriers per Lian) is a genuinely replicated order-of-magnitude statement.</li>\n</ul>\n<p><strong>Single-paper claims (plausible, not yet independently replicated):</strong></p>\n<ul>\n<li>IDLe&#39;s &quot;50× less high-fidelity data&quot; (arXiv:2412.06064) — one preprint, industry group.</li>\n<li>Hänseroth&#39;s &quot;5–15× forces, 2–4 orders energy&quot; sweep (arXiv:2511.05337) — one preprint, though consistent with the tutorial numbers.</li>\n<li>MatterSim&#39;s &quot;30 configs = 900-config scratch model&quot; — authors&#39; own benchmark on their own model; no third-party replication of that exact comparison.</li>\n<li>DPA-2&#39;s &quot;1–2 orders of magnitude&quot; — authors&#39; own; directionally confirmed by the general fine-tuning literature but the factor 100 is system-dependent.</li>\n</ul>\n<p><strong>Marketing / read with care:</strong></p>\n<ul>\n<li>MACE-MP-0&#39;s abstract phrase &quot;fine-tuned on just a handful of application-specific data points to reach ab initio accuracy&quot; (arXiv:2401.00096v3, abstract read) — directionally right for narrow tasks, but the controlled studies above show ~10²–10³ points is the honest budget for force-accurate work.</li>\n<li>ANI-1ccx&#39;s &quot;billions of times faster than CCSD(T)/CBS&quot; — true arithmetic, but the meaningful comparison is vs the ~10⁶ DFT labels it still consumed.</li>\n<li>Any &quot;up to N%&quot; headline (MatterSim&#39;s 97%) is a best-case over systems; the same paper shows 15% (not 3%) for a hard solid-electrolyte case.</li>\n<li>UMA&#39;s &quot;single model without fine-tuning ≈ specialized models&quot; (abstract read) is a direct counter-claim to the fine-tuning narrative — evidence that at 0.5B-structure pretraining scale the marginal value of fine-tuning shrinks for <em>covered</em> chemistries. Fine-tuning/Δ-correction remains necessary precisely in the out-of-coverage regime (Hammer&#39;s metal–sulfur result; OMat24&#39;s softening analysis).</li>\n</ul>\n<h2 id=\"4-openings-for-lupine\">4. Openings for Lupine</h2><ol>\n<li><strong>Beat the fine-tuned-NEB baseline on its own metric, with 10–50× fewer labels.</strong> The strongest published barrier-correction result (Lian et al., JMCA 2025) needs on the order of 10³ transition-state DFT labels per model family and still carries 0.07–0.09 eV barrier MAE. Lupine&#39;s sparse-anchor protocol (4–6 DFT images near the model-predicted saddle, per path) plus theorem-gated runtime correction targets <em>per-path</em> errors rather than amortized global ones. A head-to-head on their NASICON/LATP panel — labels-per-barrier vs barrier MAE — is the cleanest possible bake-off: fine-tuning pays the label cost upfront and amortizes over many paths of one chemistry; Lupine pays ~5 labels per path and needs no retraining when the chemistry changes. Union anchors (79% fewer DFT calls on the 29-path panel) stack multiplicatively on top.</li>\n<li><strong>Turn Hammer-style Δ-GPR from a static global correction into a theorem-gated local one.</strong> The Aarhus work (arXiv:2502.21179, arXiv:2507.18485) proves that (i) uMLIP errors are systematic and representation-learnable, and (ii) a sparse-GPR Δ-model on the model&#39;s own embeddings can correct global optimization with a small active DFT set. But their correction is a single global surface with no per-query guarantee and was demonstrated for structure optimization, not barriers. Lupine can consume the same signal (embedding-space error predictability) while gating each corrected quantity — saddle height, curvature — against Lean-formalized physical theorems, and falling back to DFT anchors when the gate fails. That is a strict superset of their scheme, aimed at the barrier/NEB regime they leave open.</li>\n<li><strong>Use multi-fidelity optimal-budget theory to make &quot;sparse&quot; anchors <em>provably</em> sparse.</strong> Zaspel&#39;s CQML and the Vinod–Zaspel data-cost benchmarking (arXiv:2410.11391) formalize exactly our allocation problem: how many evaluations at each fidelity minimize total QC cost for a target accuracy. Lupine&#39;s current heuristic (~4–6 anchors per band) is a hand-tuned point in that space; a MFML-style cost model treating {uMLIP ensemble, fine-tuned uMLIP, DFT anchor} as three fidelities could derive the anchor count per path from measured inter-fidelity correlation — and the benchmarking paper&#39;s key finding (Δ-ML-style two-level schemes win precisely when <em>few</em> high-fidelity evaluations are available) is theoretical cover for the sparse-anchor regime where Lupine operates.</li>\n<li><strong>Position against fine-tuning&#39;s documented failure modes as the &quot;don&#39;t retrain, correct at runtime&quot; alternative.</strong> Three verified failure modes are exploitable openings: fine-tuning reintroduces PES softening when the fine-tune set is relaxation-heavy (OMat24 v2 analysis); naive fine-tuning fails out-of-distribution and can inject bias (Wong et al. 2026); pretrained representations carry systematic chemistry biases (Hammer 2025). Each fine-tune also costs a GPU training run <em>per model per chemistry</em>, and with N uMLIPs in a union-anchor pool that cost multiplies — whereas Lupine&#39;s anchors are shared across models by construction. A &quot;theorem-gated runtime correction vs per-model fine-tuning&quot; comparison, including the catastrophic-forgetting and softening-reintroduction cases, would be a differentiating chapter figure: fine-tuning is the right tool for stable, high-volume single chemistry; Lupine wins for breadth, barriers, and guaranteed correctness.</li>\n</ol>\n<p><strong>Honest gaps in this subdomain:</strong> no public study yet quantifies fine-tuning data-efficiency for UMA-scale models (Meta&#39;s fine-tuning numbers are not in the accessible text — marked accordingly above); Bogojeski et al. 2020 (DTNN Δ-learning) could not be fetched (Nature redirect loop) and was dropped rather than cited; Pilania/Batra multi-fidelity co-kriging classics (hafnia, ACS AMI 2019, DOI 10.1021/acsami.9b02174) were only seen as citations, not read, and are excluded from quantitative claims. The subdomain is not thin — it is currently the hottest part of the literature — but truly <em>controlled</em> fine-tune-vs-scratch experiments remain concentrated in preprints from model authors&#39; own teams.</p>\n"}