{"id":"savings-dft-systems","title":"Systems-Level DFT Acceleration","subtitle":"LS-DFT, GPU ports, mixed precision and reduced-rank methods, SCF accelerators, k-point economization, and wavefunction reuse.","category":"references","tags":["literature-review","savings-stack","dft","systems","hpc"],"source":"articles/lit-review/savings-dft-systems.md","lang":"en","words":2852,"readMinutes":13,"toc":[{"depth":2,"text":"1. Executive result","id":"1-executive-result"},{"depth":2,"text":"2. Savings table","id":"2-savings-table"},{"depth":2,"text":"3. Proven vs claimed","id":"3-proven-vs-claimed"},{"depth":2,"text":"4. Openings for Lupine","id":"4-openings-for-lupine"}],"html":"<blockquote>\n<p><strong>Provenance:</strong> explore agent <code>agent-17</code> (director-commissioned deep research, swarm of 7, 2026-07-21) — materialized verbatim, then editorially corrected only to replace the retracted 624/132 union-anchor figures with the reproducible primary-record values (558/154) and update arithmetic derived from them. Quantitative literature claims are as reported by the research agent from sources it accessed; see citations inline. Citation-verification pass pending before any external publication.</p>\n</blockquote>\n<h1 id=\"chapter-digest-systems-level-dft-acceleration-ls-dft-gpu-ports-mixed-precision-reduced-rank-scf-accelerators-k-point-economization-wfn-density-reuse-ml-guesses\">Chapter digest: Systems-level DFT acceleration (LS-DFT, GPU ports, mixed precision / reduced-rank, SCF accelerators, k-point economization, wfn/density reuse, ML guesses)</h1><p>Evidence cut: 2026-07-21. Every number below was read this session; each citation is tagged <strong>[full text]</strong> or <strong>[abstract only]</strong>. Items I could not verify are marked [UNVERIFIED].</p>\n<h2 id=\"1-executive-result\">1. Executive result</h2><ul>\n<li><strong>GPU ports give 3–20× per-node wall-clock speedups, but only above a size threshold.</strong> Quantum ESPRESSO-GPU (1×V100 vs 18-core Xeon): 1.4× at 98 atoms/246 electrons, rising to &gt;3× as size grows [full text, <a href=\"https://ar5iv.labs.arxiv.org/html/2104.10502\">Giannozzi et al., JCP 152, 154105 (2020)</a>]. At the engineered end, DFT-FE reports ~20× CPU→GPU node speedup and 80–140 s ground states for 6k–15k-electron systems [abstract, <a href=\"https://arxiv.org/abs/2203.07820\">Das et al., arXiv:2203.07820v2</a>]; GPU-SPARC hybrid functionals: up to 8× node-hours / 80× core-hours, ~300 s for a &gt;6,000-electron metallic system [abstract, <a href=\"https://arxiv.org/abs/2501.16572\">Jing et al., arXiv:2501.16572</a>]. GPU speedups are workload- and size-dependent, not flat marketing numbers.</li>\n<li><strong>Reduced-rank exchange (ACE, ACE-ISDF, + mixed precision) is the biggest algorithmic win for hybrid DFT: ~2 orders of magnitude cheaper exact exchange.</strong> Poisson solves O(N_e²)→O(N_e); a converged HSE calculation on 1,000-atom bulk Si in 10 min on 2,000 cores [abstract, <a href=\"https://arxiv.org/abs/1707.09141\">Hu, Lin, Yang, arXiv:1707.09141</a>]; per-SCF-iteration cost only &quot;marginally larger than GGA&quot; [abstract, <a href=\"https://arxiv.org/abs/1601.07159\">Lin, ACE, arXiv:1601.07159</a>]. Dual-grid + single-precision ISDF/ACE/FFT adds a further &quot;several-fold&quot; (≈2× for the FFT-heavy part) at acceptable accuracy cost; 8,000-atom Si hybrid on 16,000 CPU cores [abstract, <a href=\"https://pubmed.ncbi.nlm.nih.gov/?term=Dual-Grid+and+Mixed-Precision+Methods+for+Accelerating+Hybrid+Density+Functional\">Hou et al., JCTC 21, 787 (2025), DOI 10.1021/acs.jctc.4c01541</a>].</li>\n<li><strong>Linear-scaling DFT changes the feasible system size, not the speed at fixed size — and it has a real crossover.</strong> CONQUEST: thousands of atoms by diagonalization, millions by LS [full text, <a href=\"https://ar5iv.labs.arxiv.org/html/2002.07704\">Nakata et al., JCP 152, 164112 (2020)</a>]; on an 8-atom cell its LS solver is ~2× <em>slower</em> than full TZTP diagonalization — LS only pays at scale. ONETEP/BigDFT/CP2K report thousands of atoms at PW accuracy with LS effort [abstracts: <a href=\"https://pubmed.ncbi.nlm.nih.gov/32384832/\">Prentice et al., DOI 10.1063/5.0004445</a>; <a href=\"https://pubmed.ncbi.nlm.nih.gov/33687268/\">Ratcliff et al., DOI 10.1063/5.0004792</a>; <a href=\"https://arxiv.org/abs/2003.03868\">Kühne et al., arXiv:2003.03868</a>].</li>\n<li><strong>SCF-iteration economics is the cheapest, most portable lever: 3× fewer iterations on hard metals, elimination of per-step SCF in AIMD, and 20–33% iteration cuts from ML initial guesses.</strong> Periodic Pulay: 3× fewer iterations than DIIS on Pd-bulk, converges at 25 K where Pulay fails [full text, <a href=\"https://ar5iv.labs.arxiv.org/html/1512.01604\">Banerjee et al., CPL 647, 31 (2016)</a>]. XL-BOMD removes the iterative ground-state optimization per MD step with no systematic energy drift [abstract, <a href=\"https://arxiv.org/abs/1705.10845\">Niklasson, JCP 147, 054103 (2017)</a>]. ML density guesses: −33.3% SCF iterations transferable to 900-atom systems [abstract, <a href=\"https://arxiv.org/abs/2509.25724\">Liu et al. (ByteDance), arXiv:2509.25724v3</a>]; −~20% <em>total</em> DFT cost for periodic materials, the first end-to-end-positive result — only because inference is 633× faster than prior density models [full text, <a href=\"https://arxiv.org/html/2601.19966v2\">Ærtebjerg et al., ELECTRAFI, arXiv:2601.19966v2</a>].</li>\n<li><strong>k-point economization is real but underused: 60% fewer k-points for metals at 1 meV/atom</strong> using generalized regular grids vs Monkhorst-Pack [abstract, <a href=\"https://arxiv.org/abs/1804.04741\">Morgan et al., Comput. Mater. Sci. 153, 424 (2018)</a>]. Smearing economization has essentially no quantified 2015–2026 literature — it&#39;s practitioner folklore (honest gap).</li>\n</ul>\n<h2 id=\"2-savings-table\">2. Savings table</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Technique</th>\n<th>What cost it removes</th>\n<th>Measured savings (system / hardware context)</th>\n<th>Accuracy cost / failure mode</th>\n<th>Citation</th>\n<th>Evidence</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Technique\">Linear-scaling DFT (CONQUEST)</td>\n<td data-label=\"What cost it removes\">O(N³) diagonalization; enables 10⁶-atom DFT</td>\n<td data-label=\"Measured savings (system / hardware context)\">Millions of atoms LS; multi-site support functions: 4 vs 17 functions/atom (4× fewer) with exponentially converging energy error; Sakurai–Sugiura eigensolve on 194,573-atom Ge/Si: 2,399 s on 6,400 K-computer nodes</td>\n<td data-label=\"Accuracy cost / failure mode\">Truncated density-matrix range; on 8-atom cell LS solver is 2× <em>slower</em> than TZTP diagonalization (crossover); support-function/basis restrictions for sparse inverse overlap</td>\n<td data-label=\"Citation\"><a href=\"https://ar5iv.labs.arxiv.org/html/2002.07704\">Nakata et al., JCP 152, 164112 (2020)</a>, arXiv:2002.07704v3 [full text]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Linear-scaling DFT (ONETEP, BigDFT, CP2K)</td>\n<td data-label=\"What cost it removes\">Same</td>\n<td data-label=\"Measured savings (system / hardware context)\">&quot;Thousands of atoms&quot; LS at PW accuracy (ONETEP); &quot;many thousands of atoms&quot; LS (BigDFT); CP2K aimed at &quot;massively-parallel and linear-scaling&quot;</td>\n<td data-label=\"Accuracy cost / failure mode\">Localization truncation; metals need finite-temperature tricks</td>\n<td data-label=\"Citation\"><a href=\"https://pubmed.ncbi.nlm.nih.gov/32384832/\">Prentice 2020, DOI 10.1063/5.0004445</a> [abstract only]; <a href=\"https://pubmed.ncbi.nlm.nih.gov/33687268/\">Ratcliff 2020, DOI 10.1063/5.0004792</a> [abstract only]; <a href=\"https://arxiv.org/abs/2003.03868\">Kühne 2020, arXiv:2003.03868</a> [full text, partial]</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">CP2K ADMM + HFX screening</td>\n<td data-label=\"What cost it removes\">Exact-exchange cost in hybrid DFT</td>\n<td data-label=\"Measured savings (system / hardware context)\">Screening O(N⁴)→O(N²); ADMM brings HFX &quot;to within a few times the cost of conventional GGA&quot;; molopt basis trim = 10× faster density mapping for liquid water; LRIGPW: 64-water box mapped with 192 atom terms instead of ~200,000 pair terms (99% distant pairs)</td>\n<td data-label=\"Accuracy cost / failure mode\">ADMM is an approximation (auxiliary basis); operator truncation radius system-dependent</td>\n<td data-label=\"Citation\"><a href=\"https://ar5iv.labs.arxiv.org/html/2003.03868\">Kühne et al., JCP 152, 194103 (2020)</a>, DOI 10.1063/5.0007045 [full text]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, Quantum ESPRESSO</td>\n<td data-label=\"What cost it removes\">CPU→GPU per-node speed</td>\n<td data-label=\"Measured savings (system / hardware context)\">1.4× at 98 atoms → &gt;3× for larger systems (1×V100 vs 18-core Xeon E5-2697); earlier CUDA-Fortran port 2–3×; numerics identical (ΔE ≤ 2·10⁻⁸ Ry)</td>\n<td data-label=\"Accuracy cost / failure mode\">Small systems don&#39;t saturate GPU; memory per card is the bottleneck</td>\n<td data-label=\"Citation\"><a href=\"https://ar5iv.labs.arxiv.org/html/2104.10502\">Giannozzi et al., JCP 152, 154105 (2020)</a>, arXiv:2104.10502v2 [full text]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, VASP (OpenACC)</td>\n<td data-label=\"What cost it removes\">Same; also fixes CUDA-C maintainability</td>\n<td data-label=\"Measured savings (system / hardware context)\">OpenACC port matches or beats the CUDA-C port; clear advantage at 8 GPUs from better scaling (VASP 5.4.4, V100/P100 era); community benchmarks 2.9–8.8× (2×V100 vs 8 CPU nodes)</td>\n<td data-label=\"Accuracy cost / failure mode\">RMM-DIIS vs Davidson coverage; tiny jobs unaccelerated; NSC slide numbers not peer-reviewed</td>\n<td data-label=\"Citation\"><a href=\"https://cug.org/proceedings/cug2018_proceedings/includes/files/pap153s2-file1.pdf\">Maintz &amp; Wetzstein, CUG 2018 proc.</a> [full text]; <a href=\"https://www.nsc.liu.se/support/past-events/VASP_workshop_2020/seminar3.pdf\">NSC workshop slides</a></td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, FHI-aims (all-electron NAO)</td>\n<td data-label=\"What cost it removes\">Hamiltonian integration, density update, forces on GPU</td>\n<td data-label=\"Measured savings (system / hardware context)\">2.4–6.6× for key steps; 3–4× full calculations over a 103-material test set; near-ideal scaling on 375-atom Bi₂Se₃</td>\n<td data-label=\"Accuracy cost / failure mode\">Only real-space ops offloaded; eigensolver (ELPA) separate</td>\n<td data-label=\"Citation\"><a href=\"https://ui.adsabs.harvard.edu/abs/2020CoPhC.25407314H/abstract\">Huhn et al., Comput. Phys. Commun. 254, 107314 (2020)</a>, DOI 10.1016/j.cpc.2020.107314 [abstract + description in arXiv:2409.09399]</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, ABACUS (LCAO)</td>\n<td data-label=\"What cost it removes\">Grid integrals + diagonalization on GPU</td>\n<td data-label=\"Measured savings (system / hardware context)\">Demonstrated up to 10,444 atoms (twisted bilayer graphene); grid integration time ~linear: 2.41 s @ 76 atoms → 93.95 s @ 2,524 atoms; bottleneck migrates to eigensolver (76% of SCF time at 1,876 atoms)</td>\n<td data-label=\"Accuracy cost / failure mode\">End-to-end speedup factors are in the results section I could not extract from HTML [UNVERIFIED: exact × number]; accuracy unchanged (same math)</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/html/2409.09399v3\">Zhang et al., arXiv:2409.09399v3</a> [full text, methods]</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, SPARC real-space (semilocal &amp; hybrid)</td>\n<td data-label=\"What cost it removes\">CPU node-hours &amp; core-hours</td>\n<td data-label=\"Measured savings (system / hardware context)\">Hybrid: up to 8× node-hours / 80× core-hours, ~300 s for &gt;6,000-electron metallic system (V100); semilocal predecessor: up to 6× node-hours / 60× core-hours, &lt;30 s for &gt;14,000 electrons</td>\n<td data-label=\"Accuracy cost / failure mode\">Kronecker-solver batching; ACE needed for exchange</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2501.16572\">Jing, Sharma, Pask, Suryanarayana, arXiv:2501.16572</a> [abstract]; semilocal 2023 numbers as quoted in arXiv:2501.16572 intro and arXiv:2409.09399 [secondary]</td>\n<td data-label=\"Evidence\">strong (hybrid) / moderate (semilocal, secondary)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPU port, DFT-FE (finite element)</td>\n<td data-label=\"What cost it removes\">All key kernels on GPU</td>\n<td data-label=\"Measured savings (system / hardware context)\">~20× CPU→GPU node speedup; 80–140 s ground states for 6k–15k-electron benchmarks; up to ~100,000 electrons</td>\n<td data-label=\"Accuracy cost / failure mode\">Higher-order FE meshes; GPU memory ceiling</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2203.07820\">Das et al., arXiv:2203.07820v2</a> (DFT-FE 1.0; CPC 2022) [abstract]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">GPAW GPU (CuPy)</td>\n<td data-label=\"What cost it removes\">Real-space kernels on GPU</td>\n<td data-label=\"Measured savings (system / hardware context)\">GPU support &quot;achieved with minor modifications of the GPAW code thanks to CuPy&quot; — no speedup numbers in abstract [UNVERIFIED: quantitative gain]</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2310.14776\">Mortensen et al., GPAW review, arXiv:2310.14776v2, DOI 10.1063/5.0182685</a> [abstract]</td>\n<td data-label=\"Evidence\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Reduced-rank exchange: ACE</td>\n<td data-label=\"What cost it removes\">Exchange-operator application &amp; update count</td>\n<td data-label=\"Measured savings (system / hardware context)\">Per-SCF-iteration cost &quot;only marginally larger than GGA&quot;; &quot;orders of magnitude speedup for Hartree-Fock-like calculations&quot;; advantageous already at tens of atoms</td>\n<td data-label=\"Accuracy cost / failure mode\">One-time ACE construction overhead; gap-independent</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/1601.07159\">Lin, JCTC 12, 2242 (2016), arXiv:1601.07159v2</a> [abstract]</td>\n<td data-label=\"Evidence\">strong (widely replicated in use)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Reduced-rank exchange: ISDF</td>\n<td data-label=\"What cost it removes\">Number of Poisson solves per exchange application</td>\n<td data-label=\"Measured savings (system / hardware context)\">O(N_e²)→O(N_e) Poisson solves; ~2 orders-of-magnitude exchange-cost reduction; 1,000-atom bulk-Si HSE converged in 10 min on 2,000 cores; scales to 8,192 cores at 4,096 atoms</td>\n<td data-label=\"Accuracy cost / failure mode\">ISDF interpolation error (controlled by rank parameter); metallic cases need care</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/1707.09141\">Hu, Lin, Yang, JCTC 13, 5420 (2017), arXiv:1707.09141v1</a> [abstract]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Mixed/single precision + dual grid in hybrid PW DFT</td>\n<td data-label=\"What cost it removes\">FP64 FFT and ISDF/ACE cost</td>\n<td data-label=\"Measured savings (system / hardware context)\">&quot;Several times&quot; further acceleration with &quot;acceptable trade-off in accuracy&quot;; single-precision FFTs alone ≈2× on FFT-bound parts; 8,000 Si-atom hybrid on 16,000 CPU cores</td>\n<td data-label=\"Accuracy cost / failure mode\">Accuracy trade-off is explicit but bounded; DP retained for non-HFX parts</td>\n<td data-label=\"Citation\"><a href=\"https://pubmed.ncbi.nlm.nih.gov/?term=Dual-Grid+and+Mixed-Precision+Methods+for+Accelerating+Hybrid+Density+Functional\">Hou et al., JCTC 21, 787–802 (2025), DOI 10.1021/acs.jctc.4c01541</a> [abstract]; 2× FFT detail via <a href=\"https://arxiv.org/html/2601.08077v1\">arXiv:2601.08077 background</a></td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Mixed precision in QC kernels (TeraChem lineage; DFT-FE)</td>\n<td data-label=\"What cost it removes\">FP64 tensor/GEMM cost</td>\n<td data-label=\"Measured savings (system / hardware context)\">TeraChem mixed SP/DP: ~2× on Kepler-era GPUs at near-DP accuracy; DFT-FE used mixed precision to reach 33% of benchmark performance on Summit</td>\n<td data-label=\"Accuracy cost / failure mode\">Legacy-hardware dependent (SP:DP throughput ratio); final iteration in DP recommended</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/pdf/2412.19322\">Mixed-precision survey, arXiv:2412.19322</a> [full-text section read]</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Periodic Pulay mixing</td>\n<td data-label=\"What cost it removes\">SCF iterations</td>\n<td data-label=\"Measured savings (system / hardware context)\">3× fewer iterations than DIIS (Pd-bulk); Fe-bulk @1000 K: 52 → 17 iterations (k=2); converges at 25 K where Pulay fails (&gt;250 iter); 2–3× over Pulay in an independent spectral code (ClusterES)</td>\n<td data-label=\"Accuracy cost / failure mode\">Extra parameter k; gains largest for metals/low T</td>\n<td data-label=\"Citation\"><a href=\"https://ar5iv.labs.arxiv.org/html/1512.01604\">Banerjee, Suryanarayana, Pask, CPL 647, 31–35 (2016), arXiv:1512.01604v2, DOI 10.1016/j.cplett.2016.01.033</a> [full text]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Predictive (trust-region) mixing</td>\n<td data-label=\"What cost it removes\">SCF iterations; robustness</td>\n<td data-label=\"Measured savings (system / hardware context)\">Polyak-step estimate from history; tested on 8 structures × 4 mixing schemes + 36 fixed cases; &quot;works well independent of candidate step&quot;</td>\n<td data-label=\"Accuracy cost / failure mode\">Few absolute iteration numbers in abstract [moderate detail]</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2104.04384\">Woods et al., arXiv:2104.04384</a> [abstract]</td>\n<td data-label=\"Evidence\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Wfn/density reuse in AIMD: XL-BOMD / CP2G ASPC</td>\n<td data-label=\"What cost it removes\">Iterative ground-state optimization per MD step</td>\n<td data-label=\"Measured savings (system / hardware context)\">XL-BOMD: &quot;without the requirement of an iterative, non-linear electronic ground state optimization prior to the force evaluations and without a systematic drift in the total energy&quot;; CP2G: &quot;substantially reduce the number of SCF iterations required per MD step&quot; (needs consistent tuning of propagation+corrector+thermostat)</td>\n<td data-label=\"Accuracy cost / failure mode\">Energy drift if mistuned; ASPC not time-reversible (high-order error)</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/1705.10845\">Niklasson, JCP 147, 054103 (2017), arXiv:1705.10845, DOI 10.1063/1.4985893</a> [abstract]; <a href=\"https://arxiv.org/abs/2601.12191\">CP2G best-practices, arXiv:2601.12191v2</a> [abstract]</td>\n<td data-label=\"Evidence\">strong (MD) / moderate (CP2G tuning guide)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Wfn/density reuse along NEB bands</td>\n<td data-label=\"What cost it removes\">First-SCF cost per image per NEB iteration</td>\n<td data-label=\"Measured savings (system / hardware context)\">[UNVERIFIED] — standard practice in VASP/QE (restart WFC per image), but I found no 2015–2026 paper quantifying the savings along a band</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation\">—</td>\n<td data-label=\"Evidence\">[UNVERIFIED]</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">k-point GR grids</td>\n<td data-label=\"What cost it removes\">Number of irreducible k-points at fixed accuracy</td>\n<td data-label=\"Measured savings (system / hardware context)\">Metals at 1 meV/atom: 60% faster than Monkhorst-Pack, 20% faster than simultaneously-commensurate grids (&gt;7,000 structures tested); finer density increments</td>\n<td data-label=\"Accuracy cost / failure mode\">Erratic convergence for metals remains; benefit mainly metals</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/1804.04741\">Morgan et al., Comput. Mater. Sci. 153, 424 (2018), arXiv:1804.04741v2</a> [abstract]</td>\n<td data-label=\"Evidence\">moderate (single group, large test set)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">Smearing economization</td>\n<td data-label=\"What cost it removes\">k-points needed for metal convergence</td>\n<td data-label=\"Measured savings (system / hardware context)\">[UNVERIFIED] — no dedicated quantified 2015–2026 study found; treated implicitly inside GR-grid and mixing papers</td>\n<td data-label=\"Accuracy cost / failure mode\">—</td>\n<td data-label=\"Citation\">—</td>\n<td data-label=\"Evidence\">[UNVERIFIED]</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ML density-matrix/density initial guess (molecules)</td>\n<td data-label=\"What cost it removes\">Early SCF iterations</td>\n<td data-label=\"Measured savings (system / hardware context)\">−33.3% SCF iterations, trained on ≤20-atom molecules, applied to 60-atom and 900-atom systems; Hamiltonian-prediction baselines +80% iterations or divergence</td>\n<td data-label=\"Accuracy cost / failure mode\">Molecular (Gaussian-basis) scope; periodic extension open</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2509.25724\">Liu et al. (ByteDance Seed), arXiv:2509.25724v3</a> [abstract]</td>\n<td data-label=\"Evidence\">moderate</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ML periodic density guess (ELECTRAFI)</td>\n<td data-label=\"What cost it removes\">Early SCF iterations + total DFT cost</td>\n<td data-label=\"Measured savings (system / hardware context)\">Up to ~20% total DFT compute reduction for crystalline materials; 633× faster inference than strongest prior density model (sub-second); slower density models negate savings</td>\n<td data-label=\"Accuracy cost / failure mode\">Savings capped by SCF-iteration share; plane-wave codes only so far</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/html/2601.19966v2\">Ærtebjerg et al., arXiv:2601.19966v2</a> [full text]</td>\n<td data-label=\"Evidence\">moderate (very new, single group)</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">ML spectral guess for BigDFT (LimitX)</td>\n<td data-label=\"What cost it removes\">Early SCF iterations at scale</td>\n<td data-label=\"Measured savings (system / hardware context)\">Chebyshev-spectrum prediction &quot;bypasses early SCF iterations in BigDFT&quot;; 2 TB protein-dimer dataset</td>\n<td data-label=\"Accuracy cost / failure mode\">No wall-clock numbers yet; project-stage</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/2606.00401\">arXiv:2606.00401</a> [abstract via search]</td>\n<td data-label=\"Evidence\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">DGDFT + PEXSI (adaptive local basis)</td>\n<td data-label=\"What cost it removes\">Degrees of freedom &amp; diagonalization</td>\n<td data-label=\"Measured savings (system / hardware context)\">~15 ALBs/atom at planewave accuracy (vs hundreds of PW coefficients); PEXSI avoids diagonalization; phosphorene ribbons to ~10,000 atoms</td>\n<td data-label=\"Accuracy cost / failure mode\">O(N²) PEXSI cost in 3D; can&#39;t exploit good initial guesses between geometry steps</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/abs/1506.08147\">Hu, Lin, Yang, JCP 143, 124110 (2015), arXiv:1506.08147</a> [abstract]</td>\n<td data-label=\"Evidence\">strong</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">PWDFT-SW (Sunway)</td>\n<td data-label=\"What cost it removes\">CPU heterogeneity exploitation</td>\n<td data-label=\"Measured savings (system / hardware context)\">64.8× speedup at 4,096 Si atoms; up to 16,384 C atoms (as described in the ABACUS GPU paper&#39;s related-work section; original report not directly accessed)</td>\n<td data-label=\"Accuracy cost / failure mode\">Sunway-specific</td>\n<td data-label=\"Citation\"><a href=\"https://arxiv.org/html/2409.09399v3\">via arXiv:2409.09399 [secondary]</a></td>\n<td data-label=\"Evidence\">claim</td>\n</tr>\n<tr>\n<td data-label=\"Technique\">PWmat (GPU-native PW, China)</td>\n<td data-label=\"What cost it removes\">CPU→GPU</td>\n<td data-label=\"Measured savings (system / hardware context)\">Vendor materials claim &quot;&gt;30× vs CPU version&quot; via mixed precision + GPU-native design; peer-reviewed basis: Jia et al., CPC 211, 8 (2017) [reference seen, abstract not read]</td>\n<td data-label=\"Accuracy cost / failure mode\">Marketing numbers; no system-size/error accounting in blog sources</td>\n<td data-label=\"Citation\"><a href=\"https://zhuanlan.zhihu.com/p/445061631\">vendor/press</a>; CPC ref via [arXiv:2501.16572 bibliography]</td>\n<td data-label=\"Evidence\">marketing</td>\n</tr>\n</tbody></table></div><h2 id=\"3-proven-vs-claimed\">3. Proven vs claimed</h2><p><strong>Replicated and trusted</strong></p>\n<ul>\n<li>GPU acceleration of PW/PAW DFT at 2–20× per node is confirmed independently across QE [full text], VASP [full text], FHI-aims, ABACUS, SPARC, DFT-FE, TeraChem-lineage — with consistent lessons: gains grow with system size, small cells (&lt;~100 atoms) can see ~1×, and GPU memory is the binding constraint. The honest metric is node-hours (SPARC&#39;s 8× node-hours vs 80× core-hours framing shows how raw &quot;×&quot; inflates if you count cores).</li>\n<li>ACE/ISDF reduced-rank exchange: implemented in multiple codes (PWDFT, SPARC, others) with follow-on papers (noncollinear spin, RT-TDDFT, mixed precision, MP2) — the ~2-order-of-magnitude exchange-cost reduction for thousand-atom hybrids is credible and replicated in use.</li>\n<li>Periodic Pulay: full-text verified 2–3× iteration reduction, independently observed in a second code (ClusterES), and adopted as default-style mixing in SPARC production runs (per the SPARC hybrid GPU paper).</li>\n<li>XL-BOMD / CP2G density-matrix propagation: decade-long literature lineage; the &quot;no per-step iterative SCF&quot; property is in the abstract of the 2017 JCP paper; the 2026 CP2G tutorial confirms it survives contact with real production settings (with tuning caveats).</li>\n</ul>\n<p><strong>Single-paper / single-group claims</strong></p>\n<ul>\n<li>ELECTRAFI&#39;s −20% end-to-end DFT cost for materials (2026 preprint, one group, ICML-style submission). The mechanism (inference cost eats SCF savings) is convincing, but it needs independent reproduction — especially since prior density models reportedly <em>increased</em> total time.</li>\n<li>ByteDance&#39;s −33.3% SCF iterations: one group, molecular Gaussian-basis only; the accompanying failure of Hamiltonian-based baselines (+80% or divergence) is a useful warning, but external validation is absent.</li>\n<li>GR k-point grids: one group (BYU), but 7,000+ structure tests and a simple, checkable algorithm; still under-adopted by major codes&#39; default workflows.</li>\n<li>Hou dual-grid/mixed-precision: one group (Lin Lin&#39;s collaboration), peer-reviewed (JCTC 2025); &quot;several times&quot; is abstract-level language — treat the precise factor as system-dependent.</li>\n</ul>\n<p><strong>Marketing / non-citable</strong></p>\n<ul>\n<li>PWmat &quot;&gt;30× GPU speedup&quot; appears only in vendor blogs/press; the peer-reviewed PWmat GPU papers (Jia et al., CPC 211, 8 (2017)) were not abstract-accessible this session → treat all PWmat numbers as marketing until the CPC paper is read.</li>\n<li>Generic &quot;GPUs give 10–100× DFT speedups&quot; blog claims (e.g., mqs.dk): not citable, contradicted by the measured 1.4–3× QE numbers at realistic sizes.</li>\n<li>PWDFT-SW 64.8×: real but Sunway-specific, and I only accessed it second-hand via the ABACUS paper.</li>\n</ul>\n<p><strong>Thin sub-areas (honest assessment)</strong>: ONETEP has no dedicated peer-reviewed GPU-port paper in 2015–2026 that I could find; k-point <em>smearing</em> economization and wavefunction reuse <em>along NEB bands</em> have no quantified literature — both are practitioner knowledge, marked [UNVERIFIED].</p>\n<h2 id=\"4-openings-for-lupine\">4. Openings for Lupine</h2><ul>\n<li><strong>Multiplicative stacking is the headline: sparse/union anchors sit on top of every technique here.</strong> GPU ports, LS-DFT, and ACE-ISDF cut the cost <em>per DFT evaluation</em>; sparse anchors (4–6 images near the saddle) cut the <em>number of evaluations</em> <del>4–8× vs a dense band, and union anchors cut it again 3.62× across models (our 154 vs 558 measurement). These are orthogonal multipliers: an ACE-ISDF-class hybrid anchor (</del>10 min for 1,000-atom Si on 2,000 cores, <a href=\"https://arxiv.org/abs/1707.09141\">Hu et al. 2017</a>) × sparse-anchor sampling makes <em>hybrid-quality barrier libraries</em> feasible where a dense hybrid NEB is not. Nobody in this literature combines reduced-rank hybrid DFT with reduced-image path sampling — that&#39;s a publishable gap.</li>\n<li><strong>Wavefunction/density reuse along the band is an unclaimed, unquantified win.</strong> The MD world formalized this (XL-BOMD removes per-step SCF entirely, <a href=\"https://arxiv.org/abs/1705.10845\">Niklasson 2017</a>); the NEB world has zero quantified studies. Because sparse anchors concentrate DFT images in a small configurational region near the saddle, image-to-image density extrapolation (ASPC-style, or ELECTRAFI-style ML density guess at each anchor, −20% total cost, <a href=\"https://arxiv.org/html/2601.19966v2\">arXiv:2601.19966</a>) should be <em>more</em> effective there than across a whole band — the images are closer together. Lupine can define, measure, and Lean-certify the error bound for density extrapolation across neighboring anchor images; this directly compounds with our 72.4% union-anchor saving.</li>\n<li><strong>Theorem gates as a robustness layer over aggressive accelerators.</strong> The measured failure modes in this subdomain are exactly where gates help: classical Pulay <em>diverges</em> on metals at low T where Periodic Pulay still converges (<a href=\"https://ar5iv.labs.arxiv.org/html/1512.01604\">Banerjee 2016</a>); mixed-precision and dual-grid introduce explicit, bounded accuracy trade-offs (<a href=\"https://pubmed.ncbi.nlm.nih.gov/?term=Dual-Grid+and+Mixed-Precision+Methods+for+Accelerating+Hybrid+Density+Functional\">Hou 2025</a>); ADMM is an approximation to the exchange energy (<a href=\"https://ar5iv.labs.arxiv.org/html/2003.03868\">Kühne 2020</a>). A theorem-gated runtime corrector can certify per-anchor convergence (residual bounds, force consistency between uMLIP prediction and DFT anchor) and only then accept the cheap settings — turning &quot;fast but fragile&quot; defaults (single-precision HFX, loose mixing) into gated defaults with verified fallback to DP/standard settings. That is a trust argument no GPU or mixing paper can make.</li>\n<li><strong>Union anchors exploit what these techniques cannot: cross-model agreement.</strong> All techniques above accelerate one model&#39;s DFT. Our measured 154 shared vs 558 naive evaluations (72.4% fewer; 3.62×) comes from uMLIP extremum agreement — a structural saving independent of hardware or algebra. The natural endpoint: uMLIP-predicted saddle region → union-shared sparse anchors → each anchor evaluated with GPU + ACE-ISDF + mixed-precision DFT, started from an ML density guess. Estimated stacking from verified numbers: 3.62× (union) × 3–8× (sparse vs dense) × ~2× (mixed precision, [Hou 2025]) × ~1.25× (density guess, [ELECTRAFI]) ≈ 27–72× total reduction in barrier-library DFT cost relative to naive dense double-precision CPU NEB — every factor individually cited above, the product itself being Lupine&#39;s contribution to demonstrate.</li>\n</ul>\n<p><strong>Sources accessed this session</strong> (all links above): 25+ primary papers — full text read for CONQUEST (arXiv:2002.07704v3), QE-exascale (arXiv:2104.10502v2), Periodic Pulay (arXiv:1512.01604v2), CP2K (arXiv:2003.03868, partial), ABACUS GPU (arXiv:2409.09399v3, methods), VASP OpenACC CUG 2018 (partial), ELECTRAFI (arXiv:2601.19966v2, partial); abstracts verified for ACE, ACE-ISDF, Hou, DFT-FE, SPARC-hybrid-GPU, ByteDance guess, GR grids, XL-BOMD, CP2G, ONETEP, BigDFT, GPAW, predictive mixing, DGDFT, LimitX, mixed-precision survey. [UNVERIFIED] items: ABACUS GPU end-to-end × (results section not extractable), NEB wavefunction-reuse savings, smearing economization, PWmat numbers (marketing), PWDFT-SW (secondary).</p>\n"}