{"id":"mlip-elastic-benchmark-preprint","title":"MLIP Elastic Benchmark Preprint","subtitle":"1×1×1 conventional-cell accuracy matches 3×3×3 at ~4× lower cost; operator failure diagnosis.","category":"validation","tags":["mlip","elasticity","benchmark","preprint","projection-law"],"source":"articles/mlip-elastic-benchmark/mlip-elastic-benchmark-preprint-2026-06-27.md","lang":"en","words":2820,"readMinutes":13,"toc":[{"depth":2,"text":"Abstract","id":"abstract"},{"depth":2,"text":"1. Introduction","id":"1-introduction"},{"depth":2,"text":"2. Methods","id":"2-methods"},{"depth":3,"text":"2.1 Benchmark arms","id":"2-1-benchmark-arms"},{"depth":3,"text":"2.2 Model selection and alias deduplication","id":"2-2-model-selection-and-alias-deduplication"},{"depth":3,"text":"2.3 Correction operators","id":"2-3-correction-operators"},{"depth":3,"text":"2.4 Cost model","id":"2-4-cost-model"},{"depth":2,"text":"3. Results","id":"3-results"},{"depth":3,"text":"3.1 Headline cost-accuracy table","id":"3-1-headline-cost-accuracy-table"},{"depth":3,"text":"3.2 Per-element accuracy","id":"3-2-per-element-accuracy"},{"depth":3,"text":"3.3 Cost ratios","id":"3-3-cost-ratios"},{"depth":3,"text":"3.4 Accuracy vs Tr2SCAN sensitivity","id":"3-4-accuracy-vs-tr2scan-sensitivity"},{"depth":2,"text":"4. Figures","id":"4-figures"},{"depth":2,"text":"5. Discussion","id":"5-discussion"},{"depth":3,"text":"Master-plan risk-register pivot","id":"master-plan-risk-register-pivot"},{"depth":3,"text":"Caveats","id":"caveats"},{"depth":2,"text":"6. Data availability","id":"6-data-availability"},{"depth":2,"text":"7. References","id":"7-references"}],"html":"<h1 id=\"mlip-elastic-benchmark-the-1-1-1-conventional-cell-matches-3-3-3-supercell-accuracy-at-4-lower-cost-for-matpes-cubic-metal-elasticity\">MLIP Elastic Benchmark: The 1×1×1 Conventional Cell Matches 3×3×3 Supercell Accuracy at ~4× Lower Cost for MatPES Cubic-Metal Elasticity</h1><p><strong>Lupine Project</strong><br><em>Email correspondence: <a href=\"mailto:alex@lupinesci.com\">alex@lupinesci.com</a></em></p>\n<hr>\n<h2 id=\"abstract\">Abstract</h2><p>We show that a single machine-learned interatomic potential (MLIP) calculation on the conventional 1×1×1 unit cell matches the elastic-constant accuracy of a 3×3×3 supercell reference at roughly one-fourth the core-hour cost. On a 16-element cubic-metal benchmark, the raw 1×1×1 TensorNet/PBE workflow achieves a mean C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding=\"application/x-tex\">_{ij}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.5978em;vertical-align:-0.2861em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3117em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mathnormal mtight\" style=\"margin-right:0.0572em;\">ij</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.2861em;\"><span></span></span></span></span></span></span></span></span></span> MAE of 14.55 GPa (95% CI [10.08, 19.72]) at 0.0134 core-hours, compared with 14.61 GPa (95% CI [10.16, 19.81]) at 0.0518 core-hours for the 3×3×3 reference — a 1×1×1 accuracy delta of −0.06 GPa, well inside the bootstrap uncertainty. The 1×1×1 cell is therefore 3.86× cheaper than the reference with no measurable loss of accuracy for this benchmark.</p>\n<p>A leave-one-out principal-component bias operator (v0.1), intended to remove residual model-form bias at negligible cost, was tested but was not beneficial on this MLIP set: it degraded mean MAE to 63.40 GPa. The dominant source of operator failure is element-to-element variation in error direction, which a single global LOO-PCA bias cannot capture and which overcorrects pathological cases such as Cr. We report that failure as a scientific finding.</p>\n<p>A second operator, <code>scalar-bulk</code> (v0.2), learns a leave-one-out scalar re-scaling of the bulk-modulus functional shift. On the Tr2SCAN-corrected sensitivity target it achieves mean MAE 14.13 GPa at 0.0134 core-hours, beating the 3-architecture ensemble (19.89 GPa) at 2.70× lower cost. On the PBE headline target, no single-model operator beats raw; the ensemble remains the accuracy winner at 11.60 GPa (95% CI [8.57, 15.04]) and 0.0362 core-hours. The scalar-bulk operator also proves cell-size independent: when fit on the 3×3×3 grid it gives mean MAE 14.14 GPa vs Tr2SCAN. The headline operational claims that survive are therefore: (1) supercell independence — the 1×1×1 cell replaces the 3×3×3 reference at ~4× lower cost with no measurable accuracy loss; and (2) on the Tr2SCAN-corrected target, a cheap single-model operator can outperform a 3-model ensemble.</p>\n<hr>\n<h2 id=\"1-introduction\">1. Introduction</h2><p>Elastic constants are a routine gate in computational materials discovery pipelines. Supercomputer labs and high-throughput projects currently pay for that gate in one of two currencies: large supercells that suppress finite-size artifacts, or ensembles of independent models that average away model-form error. Both are expensive. A 3×3×3 supercell contains 27 times as many atoms as the conventional cubic cell; a three- to five-model ensemble multiplies inference cost by the number of models. For a 16-element benchmark this can easily dominate the compute budget of a screening campaign.</p>\n<p>Recent work on MatPES foundation MLIPs has shown that, for cubic metals, elastic constants computed from a 1×1×1 conventional cell are statistically indistinguishable from those computed from a 3×3×3 supercell: the mean C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding=\"application/x-tex\">_{ij}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.5978em;vertical-align:-0.2861em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3117em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mathnormal mtight\" style=\"margin-right:0.0572em;\">ij</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.2861em;\"><span></span></span></span></span></span></span></span></span></span> MAE for Cu and Ni moves by only +0.03 GPa when the cell grows 27× in volume [1,2]. Finite-size effects are therefore not the binding error source; the residual ~11–15 GPa error is model-form error in the MLIP training data. That observation raises two operational questions: (1) if the small cell is already accurate <em>in terms of size</em>, can it replace the large cell as the default validation setting? and (2) can a cheap post-hoc correction remove enough model-form bias to make the small single model competitive with a multi-model ensemble?</p>\n<p>Here we answer both questions by comparing five computational arms on a 16-element cubic-metal benchmark and attaching a cache-warm, core-hour cost ledger to each. The strong result is the supercell-independence result. The operator story is mixed: the global LOO-PCA operator fails, but a scalar re-scaling of the bulk-modulus shift succeeds on the Tr2SCAN-corrected target while remaining honest on the PBE headline target.</p>\n<hr>\n<h2 id=\"2-methods\">2. Methods</h2><h3 id=\"2-1-benchmark-arms\">2.1 Benchmark arms</h3><p>For each of 16 cubic elements we compute the three independent elastic constants C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>11</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{11}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">11</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span>, C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>12</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{12}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">12</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span>, and C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>44</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{44}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">44</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span> under five protocols:</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Arm</th>\n<th>Label</th>\n<th>Description</th>\n<th>Models</th>\n<th>Cell</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Arm\">A</td>\n<td data-label=\"Label\">raw-1×1×1</td>\n<td data-label=\"Description\">Single best PBE model, conventional cell, no correction</td>\n<td data-label=\"Models\">1</td>\n<td data-label=\"Cell\">1×1×1</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">B</td>\n<td data-label=\"Label\">corrected-1×1×1</td>\n<td data-label=\"Description\">Single best PBE model, conventional cell, + v0.1 global LOO-PCA operator</td>\n<td data-label=\"Models\">1</td>\n<td data-label=\"Cell\">1×1×1</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">C</td>\n<td data-label=\"Label\">ref-3×3×3</td>\n<td data-label=\"Description\">Single best PBE model, 3×3×3 supercell, no correction</td>\n<td data-label=\"Models\">1</td>\n<td data-label=\"Cell\">3×3×3</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">D</td>\n<td data-label=\"Label\">ensemble-1×1×1</td>\n<td data-label=\"Description\">Mean of three distinct architectures, conventional cell</td>\n<td data-label=\"Models\">3</td>\n<td data-label=\"Cell\">1×1×1</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">E</td>\n<td data-label=\"Label\">scalar-bulk-1×1×1</td>\n<td data-label=\"Description\">Single best PBE model, conventional cell, + v0.2 scalar-bulk operator</td>\n<td data-label=\"Models\">1</td>\n<td data-label=\"Cell\">1×1×1</td>\n</tr>\n</tbody></table></div><p>The headline comparisons are: A vs C (does the small cell match the big-cell reference without correction?); A vs B (does the global operator help?); A vs E (does the scalar-bulk operator help?); and A vs D (what does the ensemble buy at what cost?). Costs are reported as CPU-equivalent core-hours on cache-warm runs.</p>\n<h3 id=\"2-2-model-selection-and-alias-deduplication\">2.2 Model selection and alias deduplication</h3><p>The MatPES 2025.2 release contains labels <code>M3GNet</code>, <code>CHGNet</code>, <code>TensorNet</code>, and <code>QET</code>. We treat <code>QET</code> and <code>TensorNet</code> as a single architecture because they resolve to the same checkpoint and return byte-identical C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding=\"application/x-tex\">_{ij}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.5978em;vertical-align:-0.2861em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3117em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mathnormal mtight\" style=\"margin-right:0.0572em;\">ij</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.2861em;\"><span></span></span></span></span></span></span></span></span></span> in every matched case [2]. The ensemble therefore contains three distinct architectures: M3GNet, CHGNet, and TensorNet.</p>\n<p>The single-model arms (A, B, C) use the lowest-MAE PBE model on the 16-element grid. Current grid data identify this as TensorNet/PBE (model identifier <code>TensorNet-PES-MatPES-PBE-2025.2</code>, also labeled QET/PBE), with a mean MAE of 13.25 GPa [2]. Headline results are reported against <code>TPBE_0K</code> targets; r2SCAN-shifted targets are used only as a sensitivity check.</p>\n<h3 id=\"2-3-correction-operators\">2.3 Correction operators</h3><p>Two operators are applied post-hoc to the 1×1×1 TensorNet/PBE prediction.</p>\n<p><strong>v0.1 global LOO-PCA (arm B).</strong> Implemented as <code>correct(raw, bias, shift)</code>. For each element:</p>\n<ul>\n<li><code>raw</code> is the predicted (C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>11</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{11}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">11</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span>, C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>12</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{12}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">12</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span>, C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>44</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{44}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">44</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span>) tensor from the 1×1×1 run.</li>\n<li><code>shift</code> is the functional shift <code>Tr2SCAN_0K − TPBE_0K</code>, taken from <code>targets_0K.json</code>.</li>\n<li><code>bias</code> is the first principal component of the centered TensorNet/PBE error matrix across the 16 elements, fitted leave-one-out: for element <em>e</em>, the bias vector is learned from the other 15 elements and then applied to <em>e</em>.</li>\n</ul>\n<p><strong>v0.2 scalar-bulk (arm E).</strong> Implemented as a leave-one-out scalar re-scaling of the bulk-modulus functional shift. For each held-out element, a single scalar <code>α</code> is fit on the other 15 elements so that the bulk modulus of <code>raw + α · (Tr2SCAN − TPBE)</code> matches the Tr2SCAN bulk modulus; that <code>α</code> is then applied to the held-out element&#39;s shift.</p>\n<p>Leave-one-out fitting is mandatory for both operators; in-sample bias fitting would leak the target. The operators are unit-tested in <code>lupine/python/lupine/operator.py</code>.</p>\n<h3 id=\"2-4-cost-model\">2.4 Cost model</h3><p>Wall-clock runtime is recorded in each per-case JSON under <code>runtime_seconds</code>. We convert to core-hours assuming single-process matcalc execution:</p>\n<span class=\"katex-display\"><span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\" display=\"block\"><semantics><mrow><mtext>core-hours</mtext><mo>=</mo><mtext>runtime_seconds</mtext><mo>×</mo><msub><mi>n</mi><mtext>cores</mtext></msub><mi mathvariant=\"normal\">/</mi><mn>3600</mn><mo separator=\"true\">,</mo></mrow><annotation encoding=\"application/x-tex\">\\text{core-hours} = \\text{runtime\\_seconds} \\times n_{\\text{cores}} / 3600,</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.6944em;\"></span><span class=\"mord text\"><span class=\"mord\">core-hours</span></span><span class=\"mspace\" style=\"margin-right:0.2778em;\"></span><span class=\"mrel\">=</span><span class=\"mspace\" style=\"margin-right:0.2778em;\"></span></span><span class=\"base\"><span class=\"strut\" style=\"height:1.0044em;vertical-align:-0.31em;\"></span><span class=\"mord text\"><span class=\"mord\">runtime_seconds</span></span><span class=\"mspace\" style=\"margin-right:0.2222em;\"></span><span class=\"mbin\">×</span><span class=\"mspace\" style=\"margin-right:0.2222em;\"></span></span><span class=\"base\"><span class=\"strut\" style=\"height:1em;vertical-align:-0.25em;\"></span><span class=\"mord\"><span class=\"mord mathnormal\">n</span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.1514em;\"><span style=\"top:-2.55em;margin-left:0em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord text mtight\"><span class=\"mord mtight\">cores</span></span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span><span class=\"mord\">/3600</span><span class=\"mpunct\">,</span></span></span></span></span>\n\n<p>with <code>n_cores = 1</code> for per-case CPU-equivalent core-hours. The 1×1×1 costs are taken from cache-warm runs after model downloads are complete, so that one-time HuggingFace cache misses do not inflate the small-cell cost.</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Arm</th>\n<th>Runtime source</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Arm\">A raw-1×1×1</td>\n<td data-label=\"Runtime source\">Cache-warm re-run of TensorNet/PBE 1×1×1</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">B corrected-1×1×1</td>\n<td data-label=\"Runtime source\">= A + negligible LOO-PCA algebra (&lt;1 s)</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">C ref-3×3×3</td>\n<td data-label=\"Runtime source\">Existing 3×3×3 16-element grid</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">D ensemble-1×1×1</td>\n<td data-label=\"Runtime source\">Cache-warm runs of M3GNet + CHGNet + TensorNet 1×1×1</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">E scalar-bulk-1×1×1</td>\n<td data-label=\"Runtime source\">= A + negligible scalar-bulk algebra (&lt;1 s)</td>\n</tr>\n</tbody></table></div><hr>\n<h2 id=\"3-results\">3. Results</h2><h3 id=\"3-1-headline-cost-accuracy-table\">3.1 Headline cost-accuracy table</h3><div class=\"table-wrap\"><table><thead><tr>\n<th>Arm</th>\n<th align=\"right\">Mean MAE C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding=\"application/x-tex\">_{ij}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.5978em;vertical-align:-0.2861em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3117em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mathnormal mtight\" style=\"margin-right:0.0572em;\">ij</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.2861em;\"><span></span></span></span></span></span></span></span></span></span> (GPa)</th>\n<th align=\"right\">Median MAE (GPa)</th>\n<th align=\"right\">95% CI (GPa)</th>\n<th align=\"right\">Core-hours (16 elem)</th>\n<th align=\"right\">Wall time (s)</th>\n<th>vs raw-1×1×1</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Arm\"><strong>A raw-1×1×1</strong></td>\n<td align=\"right\" data-label=\"Mean MAE C (GPa)\"><strong>14.55</strong></td>\n<td align=\"right\" data-label=\"Median MAE (GPa)\"><strong>13.48</strong></td>\n<td align=\"right\" data-label=\"95% CI (GPa)\"><strong>[10.08, 19.72]</strong></td>\n<td align=\"right\" data-label=\"Core-hours (16 elem)\"><strong>0.0134</strong></td>\n<td align=\"right\" data-label=\"Wall time (s)\"><strong>48.3</strong></td>\n<td data-label=\"vs raw-1×1×1\"><strong>= 1.0×</strong></td>\n</tr>\n<tr>\n<td data-label=\"Arm\">B corrected-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE C (GPa)\">63.40</td>\n<td align=\"right\" data-label=\"Median MAE (GPa)\">65.33</td>\n<td align=\"right\" data-label=\"95% CI (GPa)\">[57.07, 69.18]</td>\n<td align=\"right\" data-label=\"Core-hours (16 elem)\">0.0134</td>\n<td align=\"right\" data-label=\"Wall time (s)\">48.3</td>\n<td data-label=\"vs raw-1×1×1\">operator degrades accuracy</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">C ref-3×3×3</td>\n<td align=\"right\" data-label=\"Mean MAE C (GPa)\">14.61</td>\n<td align=\"right\" data-label=\"Median MAE (GPa)\">13.56</td>\n<td align=\"right\" data-label=\"95% CI (GPa)\">[10.16, 19.81]</td>\n<td align=\"right\" data-label=\"Core-hours (16 elem)\">0.0518</td>\n<td align=\"right\" data-label=\"Wall time (s)\">186.3</td>\n<td data-label=\"vs raw-1×1×1\">3.86× more expensive</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">D ensemble-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE C (GPa)\">11.60</td>\n<td align=\"right\" data-label=\"Median MAE (GPa)\">11.62</td>\n<td align=\"right\" data-label=\"95% CI (GPa)\">[8.57, 15.04]</td>\n<td align=\"right\" data-label=\"Core-hours (16 elem)\">0.0362</td>\n<td align=\"right\" data-label=\"Wall time (s)\">130.3</td>\n<td data-label=\"vs raw-1×1×1\">2.70× more expensive, best vs TPBE</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">E scalar-bulk-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE C (GPa)\">19.17</td>\n<td align=\"right\" data-label=\"Median MAE (GPa)\">19.01</td>\n<td align=\"right\" data-label=\"95% CI (GPa)\">[12.99, 26.87]</td>\n<td align=\"right\" data-label=\"Core-hours (16 elem)\">0.0134</td>\n<td align=\"right\" data-label=\"Wall time (s)\">48.3</td>\n<td data-label=\"vs raw-1×1×1\">= 1.0× cost, best vs Tr2SCAN</td>\n</tr>\n</tbody></table></div><p>The 1×1×1 versus 3×3×3 accuracy delta is −0.06 GPa, i.e. the small cell is slightly better by an amount well inside the 95% confidence intervals of both arms. The target uncertainty for equivalence of A and C is taken as 1.0 GPa; the observed delta is an order of magnitude smaller. Arm E has the same wall time as arm A; the scalar-bulk algebra is negligible.</p>\n<h3 id=\"3-2-per-element-accuracy\">3.2 Per-element accuracy</h3><div class=\"table-wrap\"><table><thead><tr>\n<th>Element</th>\n<th align=\"right\">raw-1×1×1 MAE (GPa)</th>\n<th align=\"right\">corrected-1×1×1 MAE (GPa)</th>\n<th align=\"right\">scalar-bulk-1×1×1 MAE (GPa)</th>\n<th align=\"right\">ref-3×3×3 MAE (GPa)</th>\n<th align=\"right\">ensemble-1×1×1 MAE (GPa)</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Element\">Ag</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">3.65</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">69.33</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">21.14</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">3.63</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">3.48</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Al</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">10.61</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">51.60</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">10.61</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">10.59</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">12.24</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Au</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">21.38</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">41.85</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">18.61</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">21.41</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">22.87</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Ca</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">2.51</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">59.84</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">2.51</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">2.53</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">2.49</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Cr</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">45.85</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">87.41</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">63.85</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">46.08</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">17.22</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Cu</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">9.72</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">76.53</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">31.54</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">9.73</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">11.98</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Fe</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">20.64</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">52.01</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">9.45</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">20.41</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">8.06</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Mo</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">13.03</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">57.32</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">5.04</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">13.16</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">12.87</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Nb</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">21.97</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">37.56</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">22.69</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">21.92</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">24.80</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Ni</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">9.16</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">70.99</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">26.40</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">9.68</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">8.71</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Pd</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">6.41</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">71.12</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">24.26</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">6.55</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">8.04</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Pt</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">18.65</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">65.74</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">17.42</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">18.80</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">17.19</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Sr</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">2.26</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">64.09</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">2.26</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">2.26</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">2.23</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Ta</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">18.61</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">65.25</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">10.02</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">18.63</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">6.06</td>\n</tr>\n<tr>\n<td data-label=\"Element\">V</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">13.94</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">78.28</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">21.51</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">13.97</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">11.26</td>\n</tr>\n<tr>\n<td data-label=\"Element\">W</td>\n<td align=\"right\" data-label=\"raw-1×1×1 MAE (GPa)\">14.34</td>\n<td align=\"right\" data-label=\"corrected-1×1×1 MAE (GPa)\">65.40</td>\n<td align=\"right\" data-label=\"scalar-bulk-1×1×1 MAE (GPa)\">19.41</td>\n<td align=\"right\" data-label=\"ref-3×3×3 MAE (GPa)\">14.41</td>\n<td align=\"right\" data-label=\"ensemble-1×1×1 MAE (GPa)\">16.11</td>\n</tr>\n</tbody></table></div><p>The v0.1 global LOO-PCA correction raises MAE on every element. The largest raw errors are Cr (45.85 GPa) and Nb (21.97 GPa); the global operator inflates both, with Cr rising to 87.41 GPa. The ensemble, by contrast, cuts Cr&#39;s MAE to 17.22 GPa and Ta&#39;s to 6.06 GPa. The v0.2 scalar-bulk operator has a different error geometry: it improves several transition metals (Fe, Mo, Ta) but raises MAE on noble/coinage FCC elements (Ag, Au, Cu, Pd, Ni) because it only re-scales the bulk-modulus shift and does not capture their off-diagonal error direction.</p>\n<h3 id=\"3-3-cost-ratios\">3.3 Cost ratios</h3><div class=\"table-wrap\"><table><thead><tr>\n<th>Comparison</th>\n<th align=\"right\">Cost ratio</th>\n<th align=\"right\">Accuracy ratio (MAE)</th>\n<th>Interpretation</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Comparison\">raw-1×1×1 vs ref-3×3×3</td>\n<td align=\"right\" data-label=\"Cost ratio\">3.86× cheaper</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">0.996× (delta −0.06 GPa)</td>\n<td data-label=\"Interpretation\">supercell-independence saving</td>\n</tr>\n<tr>\n<td data-label=\"Comparison\">scalar-bulk-1×1×1 vs ref-3×3×3</td>\n<td align=\"right\" data-label=\"Cost ratio\">3.86× cheaper</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">1.32× worse vs TPBE; 0.62× better vs Tr2SCAN</td>\n<td data-label=\"Interpretation\">same cost as raw, operator is cell-size independent</td>\n</tr>\n<tr>\n<td data-label=\"Comparison\">scalar-bulk-1×1×1 vs ensemble-1×1×1</td>\n<td align=\"right\" data-label=\"Cost ratio\">2.70× cheaper</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">1.65× worse vs TPBE; 0.71× better vs Tr2SCAN</td>\n<td data-label=\"Interpretation\">beats ensemble on Tr2SCAN target at single-model cost</td>\n</tr>\n<tr>\n<td data-label=\"Comparison\">ensemble-1×1×1 vs raw-1×1×1</td>\n<td align=\"right\" data-label=\"Cost ratio\">2.70× more expensive</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">0.797×</td>\n<td data-label=\"Interpretation\">accuracy improvement vs single-model cost</td>\n</tr>\n<tr>\n<td data-label=\"Comparison\">ensemble-1×1×1 vs ref-3×3×3</td>\n<td align=\"right\" data-label=\"Cost ratio\">1.43× cheaper</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">0.794×</td>\n<td data-label=\"Interpretation\">accuracy improvement vs supercell cost</td>\n</tr>\n<tr>\n<td data-label=\"Comparison\">corrected-1×1×1 vs raw-1×1×1</td>\n<td align=\"right\" data-label=\"Cost ratio\">1.00×</td>\n<td align=\"right\" data-label=\"Accuracy ratio (MAE)\">4.36× worse</td>\n<td data-label=\"Interpretation\">v0.1 global operator not beneficial on this set</td>\n</tr>\n</tbody></table></div><h3 id=\"3-4-accuracy-vs-tr2scan-sensitivity\">3.4 Accuracy vs Tr2SCAN sensitivity</h3><div class=\"table-wrap\"><table><thead><tr>\n<th>Arm</th>\n<th align=\"right\">Mean MAE vs Tr2SCAN (GPa)</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Arm\">raw-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE vs Tr2SCAN (GPa)\">22.55</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">ref-3×3×3</td>\n<td align=\"right\" data-label=\"Mean MAE vs Tr2SCAN (GPa)\">22.63</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">ensemble-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE vs Tr2SCAN (GPa)\">19.89</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">corrected-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE vs Tr2SCAN (GPa)\">54.28</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">scalar-bulk-1×1×1</td>\n<td align=\"right\" data-label=\"Mean MAE vs Tr2SCAN (GPa)\">14.13</td>\n</tr>\n</tbody></table></div><p>The r2SCAN sensitivity check changes the ranking. <code>scalar-bulk</code> is now best (14.13 GPa), ahead of the ensemble (19.89 GPa), raw (22.55 GPa), and ref-3×3×3 (22.63 GPa). The global LOO-PCA corrected arm remains far worse (54.28 GPa). This is the target on which the scalar-bulk operator is recommended.</p>\n<hr>\n<h2 id=\"4-figures\">4. Figures</h2><p><strong>Figure 1 — Cost-accuracy frontier.</strong> Scatter plot with core-hours on a logarithmic x-axis and mean C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding=\"application/x-tex\">_{ij}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.5978em;vertical-align:-0.2861em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3117em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mathnormal mtight\" style=\"margin-right:0.0572em;\">ij</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.2861em;\"><span></span></span></span></span></span></span></span></span></span> MAE on the y-axis. Five points correspond to arms A (raw-1×1×1), B (corrected-1×1×1), C (ref-3×3×3), D (ensemble-1×1×1), and E (scalar-bulk-1×1×1). Each point carries a vertical 95% confidence interval on MAE, computed by bootstrap over the 16 elements. Arrows annotate the A→C and A→D cost ratios; arm E shares the same x-position as A but has a different MAE. Arm B sits at the same x-position as A but with much higher MAE.</p>\n<p><strong>Figure 2 — Supercell-size independence.</strong> Per-element MAE at 1×1×1 versus 3×3×3 for the full 16-element set, with a y = x reference line. Caption reports ΔMAE = −0.06 GPa and states that finite-size effects are not the binding error source.</p>\n<p><strong>Figure 3 — Operator versus ensemble, per element.</strong> Grouped bar chart showing, for each of the 16 elements, the MAE of the raw single model, the corrected single model, the scalar-bulk single model, the 3×3×3 reference, and the ensemble mean. Bars are colored by the lowest-MAE workflow per element; the corrected single model is never the lowest.</p>\n<p><strong>Figure 4 — Error stratification by bonding class.</strong> Box plot of per-element MAE grouped into alkaline-earth FCC (Ca, Sr), noble/coinage FCC (Cu, Ag, Au), post-transition (Al), and 3d/4d/5d transition BCC+FCC (Cr, Fe, Mo, Nb, Ni, Pd, Pt, Ta, V, W).</p>\n<hr>\n<h2 id=\"5-discussion\">5. Discussion</h2><p>The central result is operational and positive: the 1×1×1 conventional cell is statistically equivalent to the 3×3×3 supercell for MatPES cubic-metal elastic constants, at roughly one-fourth the core-hour cost. The mean MAE difference is −0.06 GPa, and the 95% confidence intervals overlap substantially. For labs that currently run 3×3×3 reference calculations, switching to the 1×1×1 cell eliminates the supercell tax with no measurable accuracy penalty on this benchmark.</p>\n<p>The v0.1 global LOO-PCA correction operator was not beneficial on this MLIP set. A global bias vector is dominated by element-to-element variation in the error tensor. Cr, already the worst single element at 45.85 GPa, is driven to 87.41 GPa by the correction. The operator also produces unphysical tensor components for several elements (e.g. Sr C<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msub><mrow></mrow><mn>44</mn></msub></mrow><annotation encoding=\"application/x-tex\">_{44}</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.4511em;vertical-align:-0.15em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t vlist-t2\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.3011em;\"><span style=\"top:-2.55em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\"><span class=\"mord mtight\">44</span></span></span></span></span><span class=\"vlist-s\">​</span></span><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.15em;\"><span></span></span></span></span></span></span></span></span></span> becomes negative). The most plausible explanation is that the first principal component of the 16-element error matrix is pulled by the largest-error elements and then applied uniformly, overcorrecting the majority of elements whose error direction differs. This is a scientific finding about the limits of a global, low-rank bias correction on a small, chemically diverse set, not a failure of the underlying supercell-independence observation.</p>\n<p>The v0.2 scalar-bulk operator takes a different, more constrained approach: it learns a single scalar re-scaling of the bulk-modulus functional shift. Because the shift itself already removes much of the systematic stiffness bias, re-scaling it is a well-posed one-parameter problem. On the Tr2SCAN-corrected target it achieves mean MAE 14.13 GPa, beating the ensemble (19.89 GPa) and the raw/ref arms (~22.6 GPa) at the same single-model cost. It is also cell-size independent: when the LOO alphas are fit on the 3×3×3 grid, the mean MAE is 14.14 GPa vs Tr2SCAN. On the PBE headline target, however, scalar-bulk does not beat raw (19.17 vs 14.55 GPa) or the ensemble (11.60 GPa). The operator is therefore target-dependent: it is recommended only when the scientific goal is the Tr2SCAN-corrected elastic tensor, not the raw PBE comparison.</p>\n<p>The ensemble remains the accuracy winner on the PBE headline target. At 11.60 GPa mean MAE it improves on the single model by ~20% and remains cheaper than the 3×3×3 reference (0.0362 vs 0.0518 core-hours). For campaigns where absolute accuracy against PBE is paramount and a 2.7× cost increase over the single model is acceptable, the ensemble is the rational choice. For campaigns targeting the Tr2SCAN-corrected tensor, scalar-bulk is the rational single-model choice. For campaigns where cost is paramount, the raw 1×1×1 single model is the rational default.</p>\n<h3 id=\"master-plan-risk-register-pivot\">Master-plan risk-register pivot</h3><p>The original hypothesis had two claims: (1) supercell independence, and (2) operator-based bias removal. The first claim survives and is the primary headline. The second claim is revised: the v0.1 global LOO-PCA operator degrades accuracy, but the v0.2 scalar-bulk operator succeeds on the Tr2SCAN-corrected target while remaining honest on the PBE headline target. The 10× cost-reduction framing that relied on the operator matching the 3×3×3 reference is retired in favor of two honest claims: the ~4× cost reduction from removing the supercell, and the 2.70× cost reduction from replacing the ensemble with a scalar-bulk-corrected single model on the Tr2SCAN target. The risk-register lesson is that the supercell-independence result is robust enough to stand alone, and that operator design must be target-aware and minimally parametric.</p>\n<h3 id=\"caveats\">Caveats</h3><p>The following limitations must accompany any use of these numbers:</p>\n<ol>\n<li><p><strong>r2SCAN targets are approximated.</strong> The <code>Tr2SCAN_0K</code> tensors are PBE tensors scaled by a scalar bulk-modulus ratio [3]. This assumes shear constants scale with the bulk modulus, which is not generally true. Al, Ca, and Sr have no r2SCAN shift (<code>shift_factor = 1.0</code>). Headline numbers are reported against <code>TPBE_0K</code>; r2SCAN is a sensitivity check.</p>\n</li>\n<li><p><strong>Au uses a PW91-GGA fallback</strong>, not PBE. No stable published PBE cubic Au tensor was recovered from the de Jong 2015 dataset, AFLOW, OQMD, JARVIS-DFT, Alexandria, or the Materials Project; the PW91-GGA values of Wang &amp; Li [4] are therefore used as the reference baseline.</p>\n</li>\n<li><p><strong>QET≡TensorNet.</strong> The model roster contains three distinct architectures, not four. The QET label is an alias for TensorNet in MatPES 2025.2.</p>\n</li>\n<li><p><strong>Operator performance is target-dependent.</strong> The v0.1 global LOO-PCA operator fails on both targets. The v0.2 scalar-bulk operator is recommended only for the Tr2SCAN-corrected target; on the PBE headline target it does not beat raw or the ensemble.</p>\n</li>\n<li><p><strong>Bias is leave-one-out.</strong> Any in-sample bias fit invalidates the accuracy claim. The reported corrected-1×1×1 and scalar-bulk-1×1×1 MAEs are means over 16 LOO predictions.</p>\n</li>\n<li><p><strong>Costs are cache-warm, single-relax.</strong> Variance was checked on a four-element, three-seed subset (Ca, Cu, Fe, Cr); TensorNet/PBE 1×1×1 is deterministic and MAE standard deviations are ~0 GPa on that subset. Headline costs do not include cold-cache model downloads or seed-to-seed variability across the full set.</p>\n</li>\n</ol>\n<hr>\n<h2 id=\"6-data-availability\">6. Data availability</h2><p>The benchmark results are serialized in <code>mlip_elastic_benchmark_results.json</code> (schema <code>lupine.mlip_elastic_benchmark.v1</code>). The schema contains per-arm aggregates, per-element records, cost ratios, provenance metadata, and the caveat flags required for downstream interpretation. The raw per-case outputs, the aggregation driver, the Apptainer recipe, and a smoke-test verifier are packaged in the HPC artifact repository at <code>/home/alex/Dev/lupine/lupine-mlip-benchmark/</code>. The 16-element 3×3×3 grid and target provenance are available in the parent Lupine data store at <code>lupine/data/layer2_outputs_3x3x3_16elem/</code> and <code>lupine/data/targets_0K.json</code>.</p>\n<hr>\n<h2 id=\"7-references\">7. References</h2><p>[1] M. de Jong <em>et al.</em>, &quot;Charting the complete elastic properties of inorganic crystalline compounds,&quot; <em>Scientific Data</em> <strong>2</strong>, 150009 (2015). doi:10.1038/sdata.2015.9</p>\n<p>[2] Lupine Project, &quot;Results — Round 2: The Projection Law Correction Operator,&quot; <code>exports/library-content/latest/articles/docs/projection-law-round2-results.md</code> (2026-06-26). MatPES 2025.2; QET≡TensorNet alias deduplicated.</p>\n<p>[3] Y. Liu <em>et al.</em>, &quot;r<span class=\"katex\"><span class=\"katex-mathml\"><math xmlns=\"http://www.w3.org/1998/Math/MathML\"><semantics><mrow><msup><mrow></mrow><mn>2</mn></msup></mrow><annotation encoding=\"application/x-tex\">^2</annotation></semantics></math></span><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"strut\" style=\"height:0.8141em;\"></span><span class=\"mord\"><span></span><span class=\"msupsub\"><span class=\"vlist-t\"><span class=\"vlist-r\"><span class=\"vlist\" style=\"height:0.8141em;\"><span style=\"top:-3.063em;margin-right:0.05em;\"><span class=\"pstrut\" style=\"height:2.7em;\"></span><span class=\"sizing reset-size6 size3 mtight\"><span class=\"mord mtight\">2</span></span></span></span></span></span></span></span></span></span></span>SCAN-based DFT for materials: a benchmark and an assessment,&quot; <em>J. Chem. Phys.</em> <strong>160</strong>, 024102 (2024). doi:10.1063/5.0186586</p>\n<p>[4] L. Wang and X. Li, &quot;Ab initio calculations of elastic properties of Au at high pressure,&quot; <em>J. Appl. Phys.</em> <strong>104</strong>, 113511 (2008). doi:10.1063/1.3035832</p>\n<p>[5] T. Chen and S. P. Ong, &quot;A universal graph deep learning interatomic potential for the periodic table,&quot; <em>Nature Computational Science</em> <strong>1</strong>, 319 (2023); MatGL/MatCalc toolkit, <a href=\"https://github.com/materialsvirtuallab/matgl\">https://github.com/materialsvirtuallab/matgl</a>.</p>\n"}