{"id":"savings-techniques-synthesis-2026-07-21","title":"The Savings Stack and the Theorem Commons","subtitle":"Synthesis of the nine digests: seven layers of compute savings, each with the same hole — the missing correctness certificate.","category":"references","tags":["literature-review","savings-stack","synthesis","correctness-certificate"],"source":"articles/lit-review/savings-techniques-synthesis-2026-07-21.md","lang":"en","words":2097,"readMinutes":10,"toc":[{"depth":2,"text":"1. The argument in one paragraph","id":"1-the-argument-in-one-paragraph"},{"depth":2,"text":"2. The cost wall, briefly","id":"2-the-cost-wall-briefly"},{"depth":2,"text":"3. The seven savings layers, and the hole in each","id":"3-the-seven-savings-layers-and-the-hole-in-each"},{"depth":2,"text":"4. What we have measured ourselves","id":"4-what-we-have-measured-ourselves"},{"depth":2,"text":"5. The theorem commons","id":"5-the-theorem-commons"},{"depth":2,"text":"6. What we do not claim","id":"6-what-we-do-not-claim"},{"depth":2,"text":"7. Next measurements","id":"7-next-measurements"},{"depth":2,"text":"References","id":"references"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> director synthesis, 2026-07-21. Draws on nine deep-research digests initially materialized verbatim in this directory (<code>savings-surrogate-neb.md</code>, <code>savings-active-learning.md</code>, <code>savings-delta-multifidelity.md</code>, <code>savings-abstention-economics.md</code>, <code>savings-dft-systems.md</code>, <code>savings-electronic-surrogates.md</code>, <code>savings-path-sampling-algorithms.md</code>, <code>supercomputing-atomistic-history.md</code>, <code>research-compute-resourcing.md</code>); six digests subsequently received narrowly scoped editorial corrections replacing the retracted 624/132 union-anchor figures. It also draws on two primary records: <code>data/candidates/z1-union-anchor-economics.json</code> (+ <code>docs/analysis/z1-union-anchor-economics.md</code>) and <code>docs/plans/2026-07-20-sparse-dft-pilot-preregistration.md</code>. The identifier-level citation audit in <code>citation-verification-2026-07-21.md</code> verified every extracted arXiv ID and DOI; literature numbers remain qualified by each digest&#39;s stated access level and [UNVERIFIED] claim flags. Every number marked <strong>[derived]</strong> is our arithmetic on stated inputs, not a published value.</p>\n</blockquote>\n<h1 id=\"the-savings-stack-and-the-theorem-commons\">The savings stack and the theorem commons</h1><p><strong>Scope:</strong> compute-savings techniques in atomistic and molecular simulation, 2015–2026, organized as one argument: the field has independently discovered seven layers of savings, each layer has a known hole, and the holes are all the same shape — the absence of a <em>correctness certificate</em>. That shape is where Lupine lives, and it is why a shared, formally verified theorem library is not a philosophy but the next savings technology.</p>\n<h2 id=\"1-the-argument-in-one-paragraph\">1. The argument in one paragraph</h2><p>Atomistic simulation spent forty years making FLOPs cheaper (capex per peak PFLOP fell ~30× in a decade: Tianhe-2 ≈ &lt;!--MATH0--&gt;0.22M/PF — see <code>supercomputing-atomistic-history.md</code>) and the last ten years discovering that the real enemy is not the FLOP price but the <em>evaluation count</em>: how many times you must call the expensive oracle at all. Every successful savings technology of the ML era — active learning, surrogate NEB, Δ-learning, abstention gates — is an evaluation-count reduction with an unpriced correctness risk attached. Our measured contribution is that correctness-gated evaluation, shared across models, is itself a savings technology with a remarkable property: <strong>four independent models of guidance cost only ~10% more DFT than one</strong> (measured, §4). If that property generalizes — and the scaling curve says it should — then a commons in which teams contribute theorems and anchors back is a machine where <em>we all get faster together</em>, at no marginal cost to anyone.</p>\n<h2 id=\"2-the-cost-wall-briefly\">2. The cost wall, briefly</h2><p>The backdrop (full version in <code>supercomputing-atomistic-history.md</code>, <code>research-compute-resourcing.md</code>):</p>\n<ul>\n<li>Delivered supercomputing is cheap and getting cheaper: Frontier&#39;s electricity alone is roughly $800–1,600 per HPL-exaflop-hour <strong>[derived]</strong>; real applications sustain 3–10× less than HPL, so useful-FLOP costs are correspondingly higher.</li>\n<li>Access is not the bottleneck for a small team: ACCESS Explore grants (~400k credits) land in days; INCITE/ALCC, EuroHPC, and Director&#39;s Discretionary pools exist above that; spot GPU clouds cover bursts.</li>\n<li>The bottleneck is that serious atomistic campaigns still price out in the thousands-to-millions of oracle evaluations. A dense CI-NEB band is ~50–100 energy/force evaluations per image-set convergence accounting (our preregistration&#39;s accounting basis); a reaction network or screening panel multiplies that by hundreds of paths; an honest multi-model uncertainty check multiplies again by the model count. FLOP prices falling 30× per decade does not rescue a workload whose evaluation count grows with ambition.</li>\n</ul>\n<p>So the interesting question is never &quot;what does a FLOP cost&quot; but &quot;who lets you make fewer oracle calls, and when are they wrong.&quot;</p>\n<h2 id=\"3-the-seven-savings-layers-and-the-hole-in-each\">3. The seven savings layers, and the hole in each</h2><p>Each layer below is covered in depth by its own digest; one strongest measured number is quoted per layer, with its hole.</p>\n<ol>\n<li><strong>Surrogate-accelerated saddle search</strong> (<code>savings-surrogate-neb.md</code>). Local GP surrogates cut true evaluations by ~an order of magnitude, replicated for a decade: Koistinen/Jónsson 2017 (JCP 147, 152720); Garrido Torres 2019 (PRL 122, 156001), 5–25× fewer calls, cost <em>decoupled from image count</em>; Goswami 2025 (JCTC 21, 7935), ~10× on 500 molecular reactions. <strong>Hole:</strong> GP variance is a sampling-density signal, not an accuracy bound (stated plainly in the 2026 tutorial review); each search re-trains and then <em>discards</em> its surrogate.</li>\n<li><strong>Active learning</strong> (<code>savings-active-learning.md</code>). Uncertainty/committee-triggered labeling removes 2–4 orders of magnitude of DFT: VASP on-the-fly MLFF skips &gt;99% of first-principles steps (Jinnouchi/Kresse 2019); DP-GEN labeled 0.0044% of ~650M explored configurations (Zhang et al., PRM 3, 023804); FLARE trains reactive force fields in ~100–250 calls. <strong>Hole:</strong> every mainstream trigger is sufficient-but-not-necessary — DP-GEN&#39;s own authors concede the committee can agree and be wrong; FLARE&#39;s GP variance <em>underestimates</em> error under strong extrapolation; rare events are the documented miss.</li>\n<li><strong>Δ-learning, multi-fidelity, transfer</strong> (<code>savings-delta-multifidelity.md</code>). MatterSim fine-tuned to revPBE0-D3 water with <strong>30 high-fidelity configurations vs 900 from scratch</strong> (30×; arXiv:2405.04967); DPA-2 reports 1–2 orders less downstream data; classic Δ-ML reaches DFT enthalpies from 1–10% of the labels (Ramakrishnan 2015). <strong>Hole:</strong> fine-tuning reintroduces PES softening (OMat24 analysis), fails out-of-distribution, and costs a training run <em>per model per chemistry</em> — the wrong shape for breadth.</li>\n<li><strong>Abstention economics</strong> (<code>savings-abstention-economics.md</code>). The cleanest production measurement: AdsorbML finds the DFT-level adsorption minimum in 87.4% of ~1000 systems while paying full DFT only on the ~13% tail (Lan et al., npj Comput. Mater. 2023). <strong>Hole:</strong> no published budget prices abstention <em>across models</em> — every number is per-model/per-potential. The cross-model sharing economy is unmeasured territory.</li>\n<li><strong>Systems-level DFT acceleration</strong> (<code>savings-dft-systems.md</code>). GPU ports give 3–20× per node above a size threshold; reduced-rank exact exchange (ACE/ACE-ISDF) cuts hybrid-DFT cost ~2 orders (1,000-atom Si HSE in 10 min on 2,000 cores); Periodic Pulay gives 3× fewer SCF iterations on hard metals; ML density guesses cut ~20–33% of SCF iterations (ELECTRAFI: −20% <em>total</em> cost). <strong>Hole:</strong> these cut the cost <em>per evaluation</em>, not the count — and the aggressive settings (mixed precision, loose mixing) carry exactly the silent-failure risk that gates exist to catch.</li>\n<li><strong>Electronic-structure surrogates</strong> (<code>savings-electronic-surrogates.md</code>). ML Hamiltonians eliminate 100% of SCF at inference (DeepH: 10³× on a MoS₂ supercell; HamGNN: 4,284-atom Si Hamiltonian in 36 s); M-OFDFT reaches chemical accuracy at 27.4× on a 738-atom protein. <strong>Hole:</strong> predicted Hamiltonians do not give variationally reliable energies/forces for MD or NEB — and ML functionals fail silently (DM21&#39;s transition-metal convergence failures).</li>\n<li><strong>Path and sampling algorithms</strong> (<code>savings-path-sampling-algorithms.md</code>). Freezing/growing-string: a TS guess in ~20–90 gradient calls (Marks et al.); ART nouveau converges a DFT saddle in 50 force evaluations vs 463 unbiased; REST2 folds trpcage with 10 replicas where T-REMD needs 48; OPES explores ~10× faster than WTMetaD. <strong>Hole:</strong> cheap exploration distorts or skips the TS region (OneOPES says so explicitly), and string methods converge to second-order saddles at measured rates (2/16, 7/24 in Marks et al.) — a <em>wrong-curvature</em> failure that is, note, a violated-theorem condition.</li>\n</ol>\n<p>Seven layers, one repeated hole: <strong>no certificate.</strong> The field&#39;s triggers, variances, and committees all estimate <em>where the model is probably wrong</em>. None can say <em>where physics is definitely violated</em>.</p>\n<h2 id=\"4-what-we-have-measured-ourselves\">4. What we have measured ourselves</h2><p>All numbers in this section trace to committed records; nothing here is from chat memory.</p>\n<p><strong>Sparse anchors work on real metal.</strong> Path mp-760344, sparse barrier 0.5567 eV vs reference 0.5244 eV → <strong>32.2 meV error, WIN</strong> against the frozen ≤40 meV gate. Anchor cost on a 4-vCPU local box: 2,483–3,668 s each (~48 min mean, ~3.2 vCPU-hours, ~&lt;!--MATH1--&gt;0.06/vCPU-h <strong>[derived]</strong>). Known line item: the GPAW-vs-VASP convention offset drifts ~122 meV along the profile (tracked as theorem-line T1).</p>\n<p><strong>The sparse protocol was validated before any DFT ran.</strong> On recorded campaign artifacts, the preregistration&#39;s simulation showed sparse-protocol barrier MAE of 1.2–9.4 meV at ~7 anchors/path versus ~50–100 dense evaluations, with the saddle image located exactly on 82–86% of paths and within ±1 on 89–93% (<code>docs/plans/2026-07-20-sparse-dft-pilot-preregistration.md</code>).</p>\n<p><strong>Union anchors: the measured sharing economy</strong> (<code>data/candidates/z1-union-anchor-economics.json</code>, recomputed 2026-07-21; supersedes earlier informal figures of 624/132/79%, which were arithmetic drift and are retracted):</p>\n<ul>\n<li><strong>558 naive per-model anchors vs 154 union anchors across 29 analyzable paths → 72.4% fewer DFT evaluations (3.62×).</strong> On the 26 fully-covered paths: 520 vs 136, 73.8% (3.82×).</li>\n<li>One path (index 14, <code>mp-756912_1_1_1_0_0</code>) is unanalyzable — it failed CI-NEB convergence in all four model artifacts. The panel denominator is 29 of 30, on record.</li>\n<li>Cross-model agreement: identical predicted saddle image on 20/29 paths (all-four basis; 26/29 within ±1); both extrema within ±1 on 12/29.</li>\n<li><strong>The scaling law is the headline.</strong> Mean over model subsets as model count k goes 1→4: naive grows 139.5 → 279 → 418.5 → 558, while union grows 139.5 → 147.8 → 152 → <strong>154</strong>. Four models of guidance cost ~10% more DFT than one. Cross-model validation — the thing everyone agrees multiplies cost — is nearly free at the oracle.</li>\n</ul>\n<p><strong>The multiplicative stack [derived estimate, undemonstrated as a product].</strong> Union-sparse evaluation (154 anchors vs a 4-model dense accounting of 4 × 50–100 = 200–400 per path) is a ~38–76× reduction in evaluation count; per-anchor cost levers from layer 5 (mixed precision ~2×, ML density guess ~1.25×) compose orthogonally to ~95–190×. Each factor is individually cited; the product is ours to demonstrate — that demonstration is the pilot program now running.</p>\n<h2 id=\"5-the-theorem-commons\">5. The theorem commons</h2><p>Here is the network effect, stated as an engineering property rather than a hope.</p>\n<p><strong>Theorems are non-rival goods with zero marginal sharing cost.</strong> A Lean-checked physical-law theorem — curvature signature at a first-order saddle, energy–force consistency, symmetry constraints, boundary conditions — is contributed once and then gates every simulation for everyone, forever. Unlike training data, a theorem leaks nothing about the contributor&#39;s chemistry or commercial targets. It is the shareable residue of work a serious team does anyway: every failed simulation is a candidate theorem, and <em>formalizing your own failure</em> converts your most expensive knowledge into a permanent asset for the commons that pays you back in everyone else&#39;s theorems.</p>\n<p><strong>The demand side already exists — the literature is asking for it.</strong> Seven digests, working independently, came back with the same hole: learned triggers are sufficient-not-necessary (DP-GEN), GP variance saturates (FLARE), ensemble σ decorrelates from true error (Annevelink–Viswanathan line), fast exploration skips the TS region (OneOPES), string methods land on second-order saddles (Marks), ML functionals fail silently (DM21). Every one of those failure modes is a <em>physical-law violation detectable without learning anything</em>. A commons of formal theorems is precisely the supply for a demand the field has been circling for five years.</p>\n<p><strong>Union anchors are the evaluation-side commons, and we have already measured its economics.</strong> Within one lab, four models share anchors at 72–74% savings, sub-linearly with model count (§4). Across labs the same arithmetic applies to anchor <em>libraries</em>: a content-addressed store mapping structure-hash → DFT energy/forces lets one team&#39;s anchor validate every other team&#39;s model on the same path. Where theorems are zero-leak by construction, anchor sharing is opt-in per project — the commons has two tiers and members choose their exposure.</p>\n<p><strong>The contribution loop.</strong> Run your campaign on the commons → your model&#39;s failures are caught by existing theorems (you save evaluations immediately) → your <em>novel</em> failures get formalized as new theorems (you contribute) → every member&#39;s gates get sharper (false-accept and false-reject rates both fall) → evaluations saved compound across the network. The gate quality is a monotone function of membership. That is what &quot;we all get faster together&quot; means mechanically: <strong>the commons converts each team&#39;s worst day into everyone&#39;s permanent speedup.</strong></p>\n<h2 id=\"6-what-we-do-not-claim\">6. What we do not claim</h2><ul>\n<li>For a single isolated saddle search with a cheap oracle, local GP surrogates win outright — the honest literature says so, and so do we. Our wedge is amortized, multi-path, multi-model workloads: panels, networks, screening campaigns.</li>\n<li>For one stable, high-volume chemistry, fine-tuning is the right tool. Runtime correction is for breadth, for barriers, and for not paying a training run per model per chemistry — and it avoids fine-tuning&#39;s documented PES-softening and OOD failures by construction.</li>\n<li>Our measured basis is one 30-path barrier panel, four uMLIPs, one DFT engine (GPAW, frozen PBE/fd settings), one chemistry family. The stacking figure in §4 is a derived estimate, not a measurement. The union scaling curve is four points deep.</li>\n<li>The nine digests retain their own [UNVERIFIED] claim flags. The identifier audit verified 115 unique arXiv IDs and 48 unique DOIs and reconciled the VASP MLFF venue ambiguity: PRB 100, 014105 / arXiv:1904.12961 is the melting-point paper; PRL 122, 225701 is the distinct earlier hybrid-perovskite demonstration. Identifier verification does not upgrade abstract-only evidence or validate every quantitative claim.</li>\n</ul>\n<h2 id=\"7-next-measurements\">7. Next measurements</h2><ol>\n<li><strong>Chgnet × 23 active paths</strong> (local pilot, running): verdict against the ≤40 meV gate plus measured wall-hours — replaces the derived unit-cost figures with measured ones. Seven ≥159-atom paths deferred (<code>data/candidates/z1-sparse-dft-deferred.json</code>), verdicts PENDING, not excluded.</li>\n<li><strong>Union-anchor driver variant:</strong> evaluate the 154 unique anchors once and assemble per-model barriers from the shared pool — the direct experimental test that shared evaluation loses nothing vs per-model evaluation.</li>\n<li><strong>Gate-versus-gate pricing:</strong> theorem-gate false-accept/false-reject rates against a GP-variance gate on the same panel, in units of barrier error per DFT call — the abstention-economics baseline the literature lacks.</li>\n<li><strong>Convergence-loosening revalidation:</strong> Gamma-point vs (2,2,2) and h=0.20 vs 0.18, one path each, adopt-if-≤5 meV — a further 4–8× and ~40% per-anchor lever pending a preregistration amendment.</li>\n</ol>\n<h2 id=\"references\">References</h2><p>Primary records: <code>data/candidates/z1-union-anchor-economics.json</code>; <code>docs/analysis/z1-union-anchor-economics.md</code>; <code>tools/analysis/union_anchor_economics.py</code>; <code>docs/plans/2026-07-20-sparse-dft-pilot-preregistration.md</code>; <code>data/candidates/z1-sparse-dft-deferred.json</code>.\nDigests (this directory): the nine files listed in the provenance header. Literature citations live inside the digests with per-item access level (full text vs abstract) and [UNVERIFIED] flags.</p>\n"}