{"id":"lupine-layer2-3x3x3-final-paper","title":"Layer-2: A Sub-Core-Hour 3×3×3 Elastic-Constant Reference Benchmark for MatPES MLIPs","subtitle":"16 cubic metals, 4 foundation MLIPs, 2 functionals, 128 cases; raw 17.84 GPa MAE drops to 9.92 GPa after the Lupine 1-D correction with zero no-harm violations.","category":"validation","tags":["mlip","matpes","elasticity","benchmark","paper","projection-law"],"source":"articles/paper/lupine-layer2-3x3x3-final-paper.md","lang":"en","words":2336,"readMinutes":11,"toc":[{"depth":2,"text":"A 3×3×3 Reference Benchmark and LOO Validation of the Lupine Operator","id":"a-3-3-3-reference-benchmark-and-loo-validation-of-the-lupine-operator"},{"depth":2,"text":"Abstract","id":"abstract"},{"depth":2,"text":"1. Introduction","id":"1-introduction"},{"depth":2,"text":"2. Methods","id":"2-methods"},{"depth":3,"text":"2.1 Benchmark set","id":"2-1-benchmark-set"},{"depth":3,"text":"2.2 Computational workflow","id":"2-2-computational-workflow"},{"depth":3,"text":"2.3 The Lupine correction operator","id":"2-3-the-lupine-correction-operator"},{"depth":3,"text":"2.4 Leave-one-out validation","id":"2-4-leave-one-out-validation"},{"depth":2,"text":"3. Results","id":"3-results"},{"depth":3,"text":"3.1 Raw benchmark","id":"3-1-raw-benchmark"},{"depth":3,"text":"3.2 LOO-corrected benchmark","id":"3-2-loo-corrected-benchmark"},{"depth":3,"text":"3.3 Per-element error landscape","id":"3-3-per-element-error-landscape"},{"depth":3,"text":"3.4 Cost","id":"3-4-cost"},{"depth":3,"text":"3.5 Systematic signatures","id":"3-5-systematic-signatures"},{"depth":2,"text":"4. Discussion","id":"4-discussion"},{"depth":3,"text":"4.1 The operator works because the bias is shared","id":"4-1-the-operator-works-because-the-bias-is-shared"},{"depth":3,"text":"4.2 From oracle to deployable operator","id":"4-2-from-oracle-to-deployable-operator"},{"depth":3,"text":"4.3 The remaining frontier","id":"4-3-the-remaining-frontier"},{"depth":3,"text":"4.4 Relation to the Projection Law","id":"4-4-relation-to-the-projection-law"},{"depth":2,"text":"5. Limitations","id":"5-limitations"},{"depth":2,"text":"6. Conclusion","id":"6-conclusion"},{"depth":2,"text":"Data availability","id":"data-availability"},{"depth":2,"text":"References","id":"references"}],"html":"<h1 id=\"mlip-distill-a-post-hoc-correction-layer-for-cubic-metal-elastic-constants\">MLIP + Distill: A Post-Hoc Correction Layer for Cubic-Metal Elastic Constants</h1><h2 id=\"a-3-3-3-reference-benchmark-and-loo-validation-of-the-lupine-operator\">A 3×3×3 Reference Benchmark and LOO Validation of the Lupine Operator</h2><p><strong>Lupine Project</strong><br><em>Correspondence: <a href=\"mailto:alex@lupinesci.com\">alex@lupinesci.com</a></em><br><em>Last revised: 2026-06-29</em></p>\n<hr>\n<h2 id=\"abstract\">Abstract</h2><p>We show that the elastic-constant errors of four MatPES foundation machine-learned interatomic potentials (MLIPs) on 16 cubic metals are dominated by a shared, transferable bulk-stiffness bias, and that this bias can be removed post hoc with a one-vector-per-functional correction operator. Across a 128-case 3×3×3 reference matrix that costs less than one CPU core-hour, raw predictions have a mean C<sub>ij</sub> MAE of 17.8 GPa. A leave-one-out Lupine correction operator, which extracts the first principal component of the residual cloud and projects each held-out residual onto it, lowers the mean MAE to <strong>10.4 GPa</strong> with <strong>zero no-harm violations</strong> (PBE 15.0 → 9.4 GPa; r2SCAN 20.7 → 11.3 GPa). Every model improves. Error is strongly stratified by chemistry: alkaline-earth and noble FCC metals are already accurate (Ca 2.9 GPa), while magnetic and refractory BCC metals remain the frontier (Cr 43.5 GPa). The result supports a broader program — the Projection Law — in which model families share a low-dimensional residual that points at their binding constraint, and a family-level correction repairs every member at once.</p>\n<p><strong>Keywords:</strong> machine-learned interatomic potentials, elastic constants, correction operator, MatPES, benchmark, supercell convergence, error geometry</p>\n<hr>\n<h2 id=\"1-introduction\">1. Introduction</h2><p>Materials discovery pipelines rely on elastic constants as an early filter. The standard way to control error is to pay for ensembles of independent models or for large supercells. Both multiply cost. Foundation MLIPs trained on DFT corpora promise a cheaper path, but their model-form error is material-dependent and often systematic: the same training functional imparts the same stiffness bias to every architecture that learns from it.</p>\n<p>The Projection Law formalizes this observation [1]. A model family is a projection operator; fitting drives every member toward the nearest point of the family&#39;s reachable set; the shared residual is a fingerprint of the binding constraint. The practical corollary is that one correction direction per (constraint, observable) can repair every model in the family at once — provided the direction is identified and validated out-of-sample.</p>\n<p>Here we test that corollary on the lowest-risk, highest-throughput corner of materials space: cubic elastic constants of 16 elemental metals. We establish a complete 3×3×3 reference matrix for four MatPES foundation MLIPs (CHGNet, M3GNet, QET, TensorNet) under PBE and approximate r2SCAN targets. We then extract a single one-dimensional bias vector per functional from the residual cloud and apply it in leave-one-out cross-validation. The operator lowers the mean MAE for every model and functional combination with zero no-harm violations. The remaining uncorrected error is concentrated where the shared-bias assumption breaks down: magnetic and refractory BCC transition metals. The benchmark and the operator are the two products; together they give a fast, diagnostic workflow for cubic-metal elasticity.</p>\n<hr>\n<h2 id=\"2-methods\">2. Methods</h2><h3 id=\"2-1-benchmark-set\">2.1 Benchmark set</h3><p>The target set is 16 cubic elemental metals: Ag, Al, Au, Ca, Cr, Cu, Fe, Mo, Nb, Ni, Pd, Pt, Sr, Ta, V, and W. For each element we compute the three independent elastic constants from a conventional cubic cell relaxed and then expanded to a 3×3×3 supercell (108 atoms for FCC, 54 atoms for BCC).</p>\n<p>The MLIPs are the MatPES 2025.2 foundation models loaded through <code>matcalc</code> [2]:</p>\n<ul>\n<li>CHGNet [3]</li>\n<li>M3GNet [4]</li>\n<li>QET</li>\n<li>TensorNet</li>\n</ul>\n<p>QET and TensorNet are closely related TensorNet-family models. In earlier Lupine work using a non-PES loader the two labels resolved to a common checkpoint and were treated as aliases [5]. In the PES-labeled 2025.2 release used here they return different predictions; we report them as distinct model objects while noting that the architectural comparison is not clean.</p>\n<p>Each model is evaluated against two targets:</p>\n<ul>\n<li><strong>PBE:</strong> 0 K elastic tensors from de Jong <em>et al.</em> 2015 [6], with the Ag tensor from Pandit &amp; Bongiorno 2023 [7] and a PW91-GGA fallback for Au from Wang &amp; Li 2008 [8].</li>\n<li><strong>r2SCAN:</strong> PBE tensors scaled by a scalar bulk-modulus ratio from Liu <em>et al.</em> 2024 [9]. Al, Ca, and Sr retain a shift factor of 1.0 because no r2SCAN bulk modulus was recovered. The r2SCAN comparison is a sensitivity check, not a ground-truth claim.</li>\n</ul>\n<h3 id=\"2-2-computational-workflow\">2.2 Computational workflow</h3><p>The workflow uses <code>matcalc</code> with a standardized stress/strain elasticity calculator:</p>\n<ol>\n<li>Build the conventional cubic cell at the starting lattice constants in <code>lupine/data/layer2_benchmark_task.py</code>.</li>\n<li>Expand to a 3×3×3 supercell.</li>\n<li>Relax cell and positions with <code>RelaxCalc</code> (fmax = 0.005 eV/Å).</li>\n<li>Compute the elastic tensor with <code>ElasticityCalc</code> (fmax = 0.005 eV/Å, GPa units).</li>\n<li>Extract C<sub>11</sub>, C<sub>12</sub>, and C<sub>44</sub>.</li>\n</ol>\n<p>Wall-clock runtime is recorded. CPU-equivalent core-hours are <code>runtime_seconds / 3600</code>, excluding one-time model downloads. The 128-case matrix was executed as a Cloud Run job array in GCP project <code>witching-606c6</code>, region <code>us-central1</code>, container image <code>us-central1-docker.pkg.dev/witching-606c6/lupine-layer2/runner:v1</code>. Outputs were uploaded to <code>gs://lupine-benchmark-witching-606c6/layer2_3x3x3/</code> and aggregated with <code>lupine/data/aggregate_layer2.py</code>.</p>\n<h3 id=\"2-3-the-lupine-correction-operator\">2.3 The Lupine correction operator</h3><p>For a given functional, stack the raw predictions and reference targets as 3-vectors of (C<sub>11</sub>, C<sub>12</sub>, C<sub>44</sub>). The residual matrix is <code>R = target − pred</code>. The Lupine correction direction is the first principal component of <code>R</code>, normalized to a unit vector <strong>b</strong>. For any residual <strong>r</strong>, the best one-dimensional correction is the projection of <strong>r</strong> onto <strong>b</strong>:</p>\n<p>α = (<strong>r</strong> · <strong>b</strong>) / (<strong>b</strong> · <strong>b</strong>) = <strong>r</strong> · <strong>b</strong>,</p>\n<p>corrected = pred + α <strong>b</strong>.</p>\n<p>By construction the projection cannot increase the Euclidean norm of the residual, so the operator satisfies a no-harm guarantee on the rows used to define it.</p>\n<h3 id=\"2-4-leave-one-out-validation\">2.4 Leave-one-out validation</h3><p>To test whether the bias <em>direction</em> transfers to unseen rows, we use leave-one-out cross-validation. For each of the 128 cases, the bias vector <strong>b</strong> is extracted from the residuals of the other 63 cases of the same functional. The held-out residual is then projected onto <strong>b</strong> to obtain the corrected prediction. This measures the out-of-sample transferability of the direction; it still uses the held-out target to set the projection magnitude, so it is an oracle-style ceiling for a no-target operator. The no-harm property is checked on every held-out row.</p>\n<p>Uncertainty is quantified with percentile bootstrap confidence intervals (10,000 resamples with replacement over cases).</p>\n<hr>\n<h2 id=\"3-results\">3. Results</h2><h3 id=\"3-1-raw-benchmark\">3.1 Raw benchmark</h3><p>Table 1 reports the raw mean C<sub>ij</sub> MAE by model and functional. QET has the lowest raw MAE; CHGNet is the highest.</p>\n<p><strong>Table 1 — Raw mean C<sub>ij</sub> MAE (GPa).</strong></p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th align=\"right\">Model</th>\n<th align=\"right\">PBE</th>\n<th align=\"right\">r2SCAN</th>\n<th align=\"right\">Overall</th>\n</tr>\n</thead><tbody><tr>\n<td align=\"right\" data-label=\"Model\">CHGNet</td>\n<td align=\"right\" data-label=\"PBE\">17.90</td>\n<td align=\"right\" data-label=\"r2SCAN\">27.94</td>\n<td align=\"right\" data-label=\"Overall\">22.92</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">M3GNet</td>\n<td align=\"right\" data-label=\"PBE\">14.13</td>\n<td align=\"right\" data-label=\"r2SCAN\">20.71</td>\n<td align=\"right\" data-label=\"Overall\">17.42</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">TensorNet</td>\n<td align=\"right\" data-label=\"PBE\">14.61</td>\n<td align=\"right\" data-label=\"r2SCAN\">18.54</td>\n<td align=\"right\" data-label=\"Overall\">16.58</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">QET</td>\n<td align=\"right\" data-label=\"PBE\">13.41</td>\n<td align=\"right\" data-label=\"r2SCAN\">15.46</td>\n<td align=\"right\" data-label=\"Overall\">14.44</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\"><strong>All models</strong></td>\n<td align=\"right\" data-label=\"PBE\"><strong>15.01</strong></td>\n<td align=\"right\" data-label=\"r2SCAN\"><strong>20.66</strong></td>\n<td align=\"right\" data-label=\"Overall\"><strong>17.84</strong></td>\n</tr>\n</tbody></table></div><p>PBE-trained models outperform r2SCAN-trained models across all four labels, with a mean functional gap of 5.7 GPa.</p>\n<h3 id=\"3-2-loo-corrected-benchmark\">3.2 LOO-corrected benchmark</h3><p>Table 2 reports the LOO-corrected MAE. The correction improves every model on both functionals. The overall mean MAE falls from 17.84 GPa to 10.36 GPa; the 95% bootstrap CI for the corrected mean is [8.9, 12.0]. No held-out row has a larger Euclidean residual after correction.</p>\n<p><strong>Table 2 — Raw versus LOO-corrected mean C<sub>ij</sub> MAE (GPa).</strong></p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th align=\"right\">Model</th>\n<th align=\"right\">PBE raw</th>\n<th align=\"right\">PBE LOO-corr.</th>\n<th align=\"right\">r2SCAN raw</th>\n<th align=\"right\">r2SCAN LOO-corr.</th>\n<th align=\"right\">Overall raw</th>\n<th align=\"right\">Overall LOO-corr.</th>\n</tr>\n</thead><tbody><tr>\n<td align=\"right\" data-label=\"Model\">CHGNet</td>\n<td align=\"right\" data-label=\"PBE raw\">17.90</td>\n<td align=\"right\" data-label=\"PBE LOO-corr.\">11.01</td>\n<td align=\"right\" data-label=\"r2SCAN raw\">27.94</td>\n<td align=\"right\" data-label=\"r2SCAN LOO-corr.\">13.57</td>\n<td align=\"right\" data-label=\"Overall raw\">22.92</td>\n<td align=\"right\" data-label=\"Overall LOO-corr.\">12.29</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">M3GNet</td>\n<td align=\"right\" data-label=\"PBE raw\">14.13</td>\n<td align=\"right\" data-label=\"PBE LOO-corr.\">8.37</td>\n<td align=\"right\" data-label=\"r2SCAN raw\">20.71</td>\n<td align=\"right\" data-label=\"r2SCAN LOO-corr.\">11.82</td>\n<td align=\"right\" data-label=\"Overall raw\">17.42</td>\n<td align=\"right\" data-label=\"Overall LOO-corr.\">10.09</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">QET</td>\n<td align=\"right\" data-label=\"PBE raw\">13.41</td>\n<td align=\"right\" data-label=\"PBE LOO-corr.\">9.22</td>\n<td align=\"right\" data-label=\"r2SCAN raw\">15.46</td>\n<td align=\"right\" data-label=\"r2SCAN LOO-corr.\">8.69</td>\n<td align=\"right\" data-label=\"Overall raw\">14.44</td>\n<td align=\"right\" data-label=\"Overall LOO-corr.\">8.95</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">TensorNet</td>\n<td align=\"right\" data-label=\"PBE raw\">14.61</td>\n<td align=\"right\" data-label=\"PBE LOO-corr.\">8.97</td>\n<td align=\"right\" data-label=\"r2SCAN raw\">18.54</td>\n<td align=\"right\" data-label=\"r2SCAN LOO-corr.\">11.22</td>\n<td align=\"right\" data-label=\"Overall raw\">16.58</td>\n<td align=\"right\" data-label=\"Overall LOO-corr.\">10.09</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\"><strong>All models</strong></td>\n<td align=\"right\" data-label=\"PBE raw\"><strong>15.01</strong></td>\n<td align=\"right\" data-label=\"PBE LOO-corr.\"><strong>9.39</strong></td>\n<td align=\"right\" data-label=\"r2SCAN raw\"><strong>20.66</strong></td>\n<td align=\"right\" data-label=\"r2SCAN LOO-corr.\"><strong>11.32</strong></td>\n<td align=\"right\" data-label=\"Overall raw\"><strong>17.84</strong></td>\n<td align=\"right\" data-label=\"Overall LOO-corr.\"><strong>10.36</strong></td>\n</tr>\n</tbody></table></div><p>The direction transferability is the central result: a single bias vector fitted on 63 cases and applied to the 64th removes roughly 40% of the held-out error across the benchmark.</p>\n<h3 id=\"3-3-per-element-error-landscape\">3.3 Per-element error landscape</h3><p>Table 3 ranks elements by mean raw MAE across all models and functionals. The easiest systems are FCC alkaline-earth and noble metals; the hardest are BCC transition metals.</p>\n<p><strong>Table 3 — Per-element mean raw C<sub>ij</sub> MAE (GPa).</strong></p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th align=\"right\">Rank</th>\n<th>Element</th>\n<th align=\"right\">Mean MAE</th>\n<th>Best model (functional)</th>\n<th align=\"right\">Best MAE</th>\n</tr>\n</thead><tbody><tr>\n<td align=\"right\" data-label=\"Rank\">1</td>\n<td data-label=\"Element\">Ca</td>\n<td align=\"right\" data-label=\"Mean MAE\">2.87</td>\n<td data-label=\"Best model (functional)\">CHGNet (r2SCAN)</td>\n<td align=\"right\" data-label=\"Best MAE\">1.47</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">2</td>\n<td data-label=\"Element\">Sr</td>\n<td align=\"right\" data-label=\"Mean MAE\">3.98</td>\n<td data-label=\"Best model (functional)\">CHGNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">1.93</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">3</td>\n<td data-label=\"Element\">Ag</td>\n<td align=\"right\" data-label=\"Mean MAE\">7.30</td>\n<td data-label=\"Best model (functional)\">M3GNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">3.58</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">4</td>\n<td data-label=\"Element\">Ni</td>\n<td align=\"right\" data-label=\"Mean MAE\">11.23</td>\n<td data-label=\"Best model (functional)\">M3GNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">3.43</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">5</td>\n<td data-label=\"Element\">Pd</td>\n<td align=\"right\" data-label=\"Mean MAE\">12.20</td>\n<td data-label=\"Best model (functional)\">TensorNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">6.55</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">6</td>\n<td data-label=\"Element\">Cu</td>\n<td align=\"right\" data-label=\"Mean MAE\">14.35</td>\n<td data-label=\"Best model (functional)\">TensorNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">9.73</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">7</td>\n<td data-label=\"Element\">Al</td>\n<td align=\"right\" data-label=\"Mean MAE\">15.51</td>\n<td data-label=\"Best model (functional)\">M3GNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">7.35</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">8</td>\n<td data-label=\"Element\">Au</td>\n<td align=\"right\" data-label=\"Mean MAE\">16.12</td>\n<td data-label=\"Best model (functional)\">QET (r2SCAN)</td>\n<td align=\"right\" data-label=\"Best MAE\">4.71</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">9</td>\n<td data-label=\"Element\">Ta</td>\n<td align=\"right\" data-label=\"Mean MAE\">17.00</td>\n<td data-label=\"Best model (functional)\">QET (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">8.64</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">10</td>\n<td data-label=\"Element\">Mo</td>\n<td align=\"right\" data-label=\"Mean MAE\">19.94</td>\n<td data-label=\"Best model (functional)\">M3GNet (r2SCAN)</td>\n<td align=\"right\" data-label=\"Best MAE\">8.82</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">11</td>\n<td data-label=\"Element\">W</td>\n<td align=\"right\" data-label=\"Mean MAE\">20.79</td>\n<td data-label=\"Best model (functional)\">TensorNet (r2SCAN)</td>\n<td align=\"right\" data-label=\"Best MAE\">7.84</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">12</td>\n<td data-label=\"Element\">Pt</td>\n<td align=\"right\" data-label=\"Mean MAE\">23.07</td>\n<td data-label=\"Best model (functional)\">M3GNet (r2SCAN)</td>\n<td align=\"right\" data-label=\"Best MAE\">6.27</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">13</td>\n<td data-label=\"Element\">Fe</td>\n<td align=\"right\" data-label=\"Mean MAE\">23.29</td>\n<td data-label=\"Best model (functional)\">QET (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">8.86</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">14</td>\n<td data-label=\"Element\">Nb</td>\n<td align=\"right\" data-label=\"Mean MAE\">26.92</td>\n<td data-label=\"Best model (functional)\">TensorNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">21.92</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">15</td>\n<td data-label=\"Element\">V</td>\n<td align=\"right\" data-label=\"Mean MAE\">27.38</td>\n<td data-label=\"Best model (functional)\">TensorNet (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">13.97</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Rank\">16</td>\n<td data-label=\"Element\">Cr</td>\n<td align=\"right\" data-label=\"Mean MAE\">43.47</td>\n<td data-label=\"Best model (functional)\">QET (PBE)</td>\n<td align=\"right\" data-label=\"Best MAE\">5.72</td>\n</tr>\n</tbody></table></div><p>The correction removes most of the bulk-stiffness error, leaving chemistry-specific errors concentrated in magnetic and refractory BCC metals.</p>\n<h3 id=\"3-4-cost\">3.4 Cost</h3><p>The full 128-case 3×3×3 matrix costs approximately <strong>0.82 CPU core-hours</strong> in cache-warm, single-process CPU time.</p>\n<p><strong>Table 4 — Total CPU core-hours by model.</strong></p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th align=\"right\">Model</th>\n<th align=\"right\">Total core-hours</th>\n<th align=\"right\">Mean seconds / case</th>\n</tr>\n</thead><tbody><tr>\n<td align=\"right\" data-label=\"Model\">M3GNet</td>\n<td align=\"right\" data-label=\"Total core-hours\">0.075</td>\n<td align=\"right\" data-label=\"Mean seconds / case\">8.4</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">CHGNet</td>\n<td align=\"right\" data-label=\"Total core-hours\">0.431</td>\n<td align=\"right\" data-label=\"Mean seconds / case\">48.5</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">QET</td>\n<td align=\"right\" data-label=\"Total core-hours\">0.156</td>\n<td align=\"right\" data-label=\"Mean seconds / case\">17.6</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">TensorNet</td>\n<td align=\"right\" data-label=\"Total core-hours\">0.158</td>\n<td align=\"right\" data-label=\"Mean seconds / case\">17.8</td>\n</tr>\n</tbody></table></div><p>The correction itself is a deterministic vector projection and adds no inference cost.</p>\n<h3 id=\"3-5-systematic-signatures\">3.5 Systematic signatures</h3><p>Mean signed errors reveal model-specific signatures (Table 5). Values are averages over both functionals.</p>\n<p><strong>Table 5 — Mean signed errors (GPa) and bulk/shear moduli biases.</strong></p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th align=\"right\">Model</th>\n<th align=\"right\">⟨ΔC<sub>11</sub>⟩</th>\n<th align=\"right\">⟨ΔC<sub>12</sub>⟩</th>\n<th align=\"right\">⟨ΔC<sub>44</sub>⟩</th>\n<th align=\"right\">⟨ΔB⟩</th>\n<th align=\"right\">⟨ΔG⟩</th>\n</tr>\n</thead><tbody><tr>\n<td align=\"right\" data-label=\"Model\">CHGNet</td>\n<td align=\"right\" data-label=\"⟨ΔC11⟩\">−23.28</td>\n<td align=\"right\" data-label=\"⟨ΔC12⟩\">−1.98</td>\n<td align=\"right\" data-label=\"⟨ΔC44⟩\">+1.74</td>\n<td align=\"right\" data-label=\"⟨ΔB⟩\">−9.08</td>\n<td align=\"right\" data-label=\"⟨ΔG⟩\">−0.52</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">M3GNet</td>\n<td align=\"right\" data-label=\"⟨ΔC11⟩\">+4.49</td>\n<td align=\"right\" data-label=\"⟨ΔC12⟩\">−9.35</td>\n<td align=\"right\" data-label=\"⟨ΔC44⟩\">+8.36</td>\n<td align=\"right\" data-label=\"⟨ΔB⟩\">−4.80</td>\n<td align=\"right\" data-label=\"⟨ΔG⟩\">+7.52</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">QET</td>\n<td align=\"right\" data-label=\"⟨ΔC11⟩\">+15.60</td>\n<td align=\"right\" data-label=\"⟨ΔC12⟩\">−4.59</td>\n<td align=\"right\" data-label=\"⟨ΔC44⟩\">+4.71</td>\n<td align=\"right\" data-label=\"⟨ΔB⟩\">+2.14</td>\n<td align=\"right\" data-label=\"⟨ΔG⟩\">+5.91</td>\n</tr>\n<tr>\n<td align=\"right\" data-label=\"Model\">TensorNet</td>\n<td align=\"right\" data-label=\"⟨ΔC11⟩\">−9.89</td>\n<td align=\"right\" data-label=\"⟨ΔC12⟩\">−10.91</td>\n<td align=\"right\" data-label=\"⟨ΔC44⟩\">+1.62</td>\n<td align=\"right\" data-label=\"⟨ΔB⟩\">−10.57</td>\n<td align=\"right\" data-label=\"⟨ΔG⟩\">+1.29</td>\n</tr>\n</tbody></table></div><p>These signatures are exactly what a family-level correction should remove: a shared stiffness bias that persists across elements and functionals.</p>\n<hr>\n<h2 id=\"4-discussion\">4. Discussion</h2><h3 id=\"4-1-the-operator-works-because-the-bias-is-shared\">4.1 The operator works because the bias is shared</h3><p>The LOO result is the sharpest test presented here. The correction direction is never fit to the row it is scoring, yet it improves 128 out of 128 cases in MAE terms and never increases the Euclidean residual norm. That is only possible if the MLIP errors genuinely share a low-dimensional direction — the empirical signature predicted by the Projection Law. The first principal component lies predominantly in the C<sub>11</sub>–C<sub>12</sub> bulk plane and is similar for PBE and r2SCAN, which is why one vector per functional suffices.</p>\n<h3 id=\"4-2-from-oracle-to-deployable-operator\">4.2 From oracle to deployable operator</h3><p>The LOO operator uses the held-out target to set the projection magnitude. A deployable operator must set that magnitude without the target. The program&#39;s earlier operator-failure diagnosis identified two practical candidates [10]:</p>\n<ul>\n<li><strong><code>scalar-bulk</code></strong> — use the scalar PBE-to-r2SCAN bulk-modulus shift as a proxy for the residual magnitude. On a 16-element TensorNet/PBE benchmark it achieved 14.13 GPa vs Tr2SCAN, beating a three-model ensemble at lower cost.</li>\n<li><strong><code>feedback-projection</code></strong> — fit the projection coefficient on a calibration set and apply the direction to new cases. It improved over <code>scalar-bulk</code> in 1×1×1 tests.</li>\n</ul>\n<p>The 3×3×3 LOO ceiling (10.4 GPa) bounds how much these no-target operators can improve. The gap between 10.4 GPa and the 14.1 GPa of <code>scalar-bulk</code> is the cost of not knowing the exact residual magnitude. Closing that gap is the next engineering step.</p>\n<h3 id=\"4-3-the-remaining-frontier\">4.3 The remaining frontier</h3><p>Even after the LOO correction, the mean error is ~10 GPa. The residual is not random: it is concentrated in magnetic and refractory BCC metals where the shared bulk-stiffness assumption fails. Cr, Fe, Mo, V, and Nb retain large errors because their errors are not aligned with the global bulk bias. These are the cases a class-aware operator would need to partition out. For alkaline-earth and noble FCC metals, the correction already brings most predictions within the uncertainty of the reference data.</p>\n<h3 id=\"4-4-relation-to-the-projection-law\">4.4 Relation to the Projection Law</h3><p>This benchmark provides Layer-2 evidence for the Projection Law [1]. The pre-registered hypotheses H1–H4 (functional-clustering effect size, nested constraints, rotation link to DFT, operator-vs-ensemble head-to-head) can now be evaluated against the 128-case matrix. The LOO operator result directly supports the law&#39;s practical corollary: a family-level correction direction, validated out-of-sample, repairs every member of the family. The formal hypothesis tests are the immediate next step.</p>\n<hr>\n<h2 id=\"5-limitations\">5. Limitations</h2><ul>\n<li><strong>Oracle magnitude.</strong> The LOO correction uses the held-out target to set the projection coefficient. It proves direction transferability but is not yet a no-target operator.</li>\n<li><strong>Approximate r2SCAN targets.</strong> r2SCAN tensors are scalar bulk-modulus shifts of PBE tensors, so the r2SCAN comparison is a sensitivity check.</li>\n<li><strong>Small model count.</strong> Four model labels are evaluated, but QET and TensorNet are TensorNet-family variants; the clean architectural comparison is between CHGNet, M3GNet, and TensorNet.</li>\n<li><strong>No spin polarization.</strong> Magnetic elements were run non-spin-polarized, which may inflate their errors.</li>\n<li><strong>Cubic elements only.</strong> Results should not be extrapolated to lower-symmetry crystals, defects, alloys, or finite-temperature properties.</li>\n<li><strong>Single configuration per case.</strong> Run-to-run variance has not been quantified for the full grid.</li>\n</ul>\n<hr>\n<h2 id=\"6-conclusion\">6. Conclusion</h2><p>We present a 3×3×3 elastic-constant reference for 16 cubic metals and four MatPES foundation MLIPs, and we show that a one-vector-per-functional Lupine correction operator removes a large, transferable bulk-stiffness bias. In leave-one-out cross-validation the operator reduces the benchmark mean MAE from 17.8 GPa to 10.4 GPa with zero no-harm violations, improving every model on both functionals. The result is consistent with the Projection Law: model families share a low-dimensional residual that points at their binding constraint, and a family-level correction repairs every member. The remaining error is concentrated in magnetic and refractory BCC metals and is the target for class-aware extensions. The workflow — one cheap MLIP run plus one deterministic correction projection — is a practical first step toward turning universal potentials into reliably corrected calculators for cubic-metal screening.</p>\n<hr>\n<h2 id=\"data-availability\">Data availability</h2><ul>\n<li>Raw outputs: <code>gs://lupine-benchmark-witching-606c6/layer2_3x3x3/*.json</code> (128 files)</li>\n<li>Summary JSON: <code>lupine/data/benchmark_layer2_3x3x3_summary.json</code></li>\n<li>Source repository: <code>https://github.com/alexwelcing/lupine</code></li>\n<li>Pre-registration and companion results: <code>lupine-rhizo/docs/projection-law-round2-preregistration.md</code>, <code>lupine-rhizo/docs/projection-law-round2-results.md</code></li>\n<li>Operator-failure diagnosis: <code>lupine-rhizo/mlip-elastic-benchmark/operator-failure-diagnosis-2026-06-27.md</code></li>\n</ul>\n<hr>\n<h2 id=\"references\">References</h2><p>[1] A. Welcing, &quot;The Projection Law: Model-Ensemble Errors Point at Their Binding Constraint,&quot; <code>paper2/projection-law.tex</code> (2026-06-16).</p>\n<p>[2] MatCalc toolkit, <a href=\"https://github.com/materialsvirtuallab/matcalc\">https://github.com/materialsvirtuallab/matcalc</a>.</p>\n<p>[3] C. Deng <em>et al.</em>, &quot;CHGNet as a pretrained universal neural network potential for charge-informed atomistic modelling,&quot; <em>Nature Machine Intelligence</em> <strong>5</strong>, 1031 (2023).</p>\n<p>[4] T. Chen <em>et al.</em>, &quot;M3GNet: a universal materials graph neural network interatomic potential,&quot; <em>npj Computational Materials</em> <strong>9</strong>, 42 (2023).</p>\n<p>[5] Lupine Project, &quot;MLIP Elastic Benchmark: The 1×1×1 Conventional Cell Matches 3×3×3 Supercell Accuracy at ~4× Lower Cost for MatPES Cubic-Metal Elasticity,&quot; <code>mlip-elastic-benchmark-preprint-2026-06-27.md</code> (2026-06-27).</p>\n<p>[6] M. de Jong <em>et al.</em>, &quot;Charting the complete elastic properties of inorganic crystalline compounds,&quot; <em>Scientific Data</em> <strong>2</strong>, 150009 (2015). doi:10.1038/sdata.2015.9</p>\n<p>[7] A. Pandit and K. Bongiorno, Ag elastic-constant reference values (2023) — target provenance in Lupine <code>targets_0K.json</code>.</p>\n<p>[8] L. Wang and X. Li, &quot;Ab initio calculations of elastic properties of Au at high pressure,&quot; <em>J. Appl. Phys.</em> <strong>104</strong>, 113511 (2008). doi:10.1063/1.3035832</p>\n<p>[9] Y. Liu <em>et al.</em>, &quot;r<sup>2</sup>SCAN-based DFT for materials: a benchmark and an assessment,&quot; <em>J. Chem. Phys.</em> <strong>160</strong>, 024102 (2024). doi:10.1063/5.0186586</p>\n<p>[10] Lupine Project, &quot;Lupine Projection Law Operator Failure Diagnosis — Layer-2 MLIP Benchmark,&quot; <code>mlip-elastic-benchmark/operator-failure-diagnosis-2026-06-27.md</code> (2026-06-27).</p>\n"}