{"id":"projection-law-round2-preregistration","title":"Round 2 Preregistration — MatPES MLIP Elastic Constants","subtitle":"Pre-registered protocol: 16 cubic elements, four MatPES potentials, two functionals, and the operator-correction kill conditions.","category":"validation","tags":["mlip","matpes","elasticity","preregistration","projection-law"],"source":"articles/docs/projection-law-round2-preregistration.md","lang":"en","words":1271,"readMinutes":6,"toc":[{"depth":2,"text":"Revision history","id":"revision-history"},{"depth":2,"text":"Systems","id":"systems"},{"depth":2,"text":"Reference targets","id":"reference-targets"},{"depth":3,"text":"Provenance summary per element","id":"provenance-summary-per-element"},{"depth":2,"text":"Model grid","id":"model-grid"},{"depth":2,"text":"Registered hypotheses and kill conditions","id":"registered-hypotheses-and-kill-conditions"},{"depth":3,"text":"H1 — Cleaned effect size","id":"h1-cleaned-effect-size"},{"depth":3,"text":"H2 — Nested constraint hierarchy","id":"h2-nested-constraint-hierarchy"},{"depth":3,"text":"H3 — Rotation link to Layer 3","id":"h3-rotation-link-to-layer-3"},{"depth":3,"text":"H4 — Compute-budget head-to-head","id":"h4-compute-budget-head-to-head"},{"depth":2,"text":"Analysis plan","id":"analysis-plan"},{"depth":2,"text":"Software artifacts","id":"software-artifacts"},{"depth":2,"text":"Scientific-integrity policy","id":"scientific-integrity-policy"}],"html":"<h1 id=\"round-2-pre-registration-protocol-the-projection-law-correction-operator\">Round 2 Pre-Registration Protocol — The Projection Law Correction Operator</h1><blockquote>\n<p><strong>Version:</strong> 2.2 (revised 2026-06-25)<br><strong>Objective:</strong> Definitively test the Projection Law&#39;s conservation-rotation mechanism by (1) removing reference-standard confounds through 0K DFT targets, (2) demonstrating operational superiority over ensemble-based UQ, and (3) producing a drop-in LAMMPS extension that any HPC user can adopt.</p>\n</blockquote>\n<h2 id=\"revision-history\">Revision history</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Date</th>\n<th>Change</th>\n<th>Author</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Date\">2026-06-25</td>\n<td data-label=\"Change\">v2.2 — Added Ag and Au targets from published literature (Ag PBE Pandit &amp; Bongiorno 2023; Au PW91-GGA Wang &amp; Li 2008 as fallback). Target set now 16 cubic elemental metals. Updated r2SCAN bulk-shift coverage and evaluated-subset note.</td>\n<td data-label=\"Author\">researcher</td>\n</tr>\n<tr>\n<td data-label=\"Date\">2026-06-26</td>\n<td data-label=\"Change\">v2.1 — Corrected PBE extraction method (direct from de Jong 2015, not matminer); updated r2SCAN fallback list to Al, Ca, Sr; added evaluated-subset note; recorded provenance for all 14 target elements.</td>\n<td data-label=\"Author\">researcher</td>\n</tr>\n</tbody></table></div><h2 id=\"systems\">Systems</h2><p>The target set comprises <strong>16 cubic elemental metals</strong> for which a published 0K elastic tensor is available:</p>\n<ul>\n<li><strong>FCC:</strong> Al, Ag, Au, Ca, Cu, Ni, Pd, Pt, Sr</li>\n<li><strong>BCC:</strong> Cr, Fe, Mo, Nb, Ta, V, W</li>\n</ul>\n<p><em>Pb, which was part of the original IMMI set, is absent from the published cubic-elastic compilations and is not in the current target set.</em></p>\n<p><strong>Evaluated subset (reported in parent benchmark):</strong> Ag, Au, Cu, Fe, Ni, Pt, V, W (8 elements). Hypotheses below are registered for the full 16-element set; where the initial benchmark covers a restricted subset, this is noted explicitly.</p>\n<h2 id=\"reference-targets\">Reference targets</h2><p>Pristine 0K elastic constants (C11, C12, C44) from:</p>\n<ul>\n<li><strong>PBE baseline:</strong> de Jong <em>et al.</em>, <em>Scientific Data</em> <strong>2</strong>, 150009 (2015). This is the published DFT elastic-tensor dataset underlying the Materials Project elasticity workflow: VASP/PBE, stress-strain finite-difference method of Le Page &amp; Saxe (Phys. Rev. B 65, 104104), 0 K static calculations. Values are extracted <strong>directly from the de Jong 2015 publication data</strong> (via the <code>matminer</code> <code>elastic_tensor_2015</code> dataset as a cross-check, but the primary source is the paper&#39;s tabulated values). The resulting file is <code>data/pbe_targets_dejong2015.json</code> (14 elements; missing Ag, Au).</li>\n<li><strong>Ag (PBE):</strong> Pandit &amp; Bongiorno, <em>Modelling Simul. Mater. Sci. Eng.</em> <strong>31</strong>, 055005 (2023). PBE elastic constants computed for a 108-atom FCC supercell with VASP, using a 480 eV plane-wave cutoff and 36×36×36 <em>k</em>-point grid. Values: C11 = 107.0, C12 = 79.0, C44 = 42.0 GPa.</li>\n<li><strong>Au (PW91-GGA fallback):</strong> Wang &amp; Li, <em>J. Phys.: Condens. Matter</em> <strong>20</strong>, 045214 (2008). The published PBE elastic-tensor entries for Au in the de Jong 2015 / Materials Project set are unphysical (negative or near-zero C44), and no stable PBE-only cubic Au tensor was recovered from AFLOW, OQMD, JARVIS-DFT, or Alexandria. We therefore adopt the PW91-GGA values from Wang &amp; Li as the reference baseline: C11 = 165.9, C12 = 142.2, C44 = 26.7 GPa. The operator benchmark is run against the r2SCAN target derived from this baseline.</li>\n<li><strong>r2SCAN:</strong> No published full r2SCAN elastic-tensor table exists for all target metals. We therefore apply a <strong>scalar bulk-modulus shift</strong> to the baseline tensors using r2SCAN/baseline bulk-modulus ratios from Liu <em>et al.</em>, <em>J. Chem. Phys.</em> <strong>160</strong>, 024102 (2024). The shift is computed as <code>Cij_r2SCAN = Cij_baseline × (B_r2SCAN / B_baseline)</code>; this preserves the tensor anisotropy (C11/C12 ratio, Zener ratio) while shifting overall stiffness. <strong>Al, Ca, and Sr</strong> lack r2SCAN bulk data in Liu <em>et al.</em> and retain the unshifted PBE baseline. The resulting file is <code>data/targets_0K.json</code> (schema <code>lupine.targets_0K.v3</code>).</li>\n</ul>\n<p>Each target value carries a provenance record (<code>material_id</code>, source citation, URL, stability flag, and fallback reason where applicable).</p>\n<h3 id=\"provenance-summary-per-element\">Provenance summary per element</h3><div class=\"table-wrap\"><table><thead><tr>\n<th>Element</th>\n<th>Baseline source</th>\n<th>r2SCAN method</th>\n<th>r2SCAN fallback reason</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Element\">Al, Ca, Sr</td>\n<td data-label=\"Baseline source\">de Jong 2015</td>\n<td data-label=\"r2SCAN method\">PBE baseline retained</td>\n<td data-label=\"r2SCAN fallback reason\">No published r2SCAN bulk modulus available</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Ag</td>\n<td data-label=\"Baseline source\">Pandit &amp; Bongiorno 2023 (PBE)</td>\n<td data-label=\"r2SCAN method\">Scalar bulk shift (Liu 2024)</td>\n<td data-label=\"r2SCAN fallback reason\">—</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Au</td>\n<td data-label=\"Baseline source\">Wang &amp; Li 2008 (PW91-GGA)</td>\n<td data-label=\"r2SCAN method\">Scalar bulk shift (Liu 2024)</td>\n<td data-label=\"r2SCAN fallback reason\">No stable published PBE Au tensor found</td>\n</tr>\n<tr>\n<td data-label=\"Element\">Cr, Cu, Fe, Mo, Nb, Ni, Pd, Pt, Ta, V, W</td>\n<td data-label=\"Baseline source\">de Jong 2015</td>\n<td data-label=\"r2SCAN method\">Scalar bulk shift (Liu 2024)</td>\n<td data-label=\"r2SCAN fallback reason\">—</td>\n</tr>\n</tbody></table></div><h2 id=\"model-grid\">Model grid</h2><p>Layer 1 — Classical interatomic potentials (OpenKIM/NIST):</p>\n<ul>\n<li>2–3 EAM potentials per element (e.g., Ackland-1987 for Cu, V, W; Ackland-1997 for Fe; Adams-1989 for Pt). The initial benchmark used the potentials available in the local NIST catalog.</li>\n</ul>\n<p>Layer 2 — Foundation MLIPs evaluated at 0K via LAMMPS plugins (registered plan):</p>\n<ul>\n<li><strong>PBE ensemble:</strong> M3GNet, CHGNet, TensorNet, QET (MatPES PBE).</li>\n<li><strong>r2SCAN ensemble:</strong> same architectures on MatPES r2SCAN.</li>\n</ul>\n<p><em>Note: The interim benchmark reported here used Layer 1 classical potentials on the 8-element subset Ag, Au, Cu, Fe, Ni, Pt, V, W. Layer 2 MLIP evaluation is staged for the full 16-element set.</em></p>\n<h2 id=\"registered-hypotheses-and-kill-conditions\">Registered hypotheses and kill conditions</h2><h3 id=\"h1-cleaned-effect-size\">H1 — Cleaned effect size</h3><p>When evaluated against 0K all-electron references, the functional-clustering effect size for 3d/4d metals meets or exceeds the Round 1 registered threshold (0.30).</p>\n<ul>\n<li><strong>Kill condition:</strong> Effect size &lt; 0.20 on the 3d/4d subset against 0K references.</li>\n<li><strong>Evaluated-subset note:</strong> The interim 8-element benchmark contains four 3d/4d metals (Cu, Fe, Ni, V). The full 16-element test will include additional 3d/4d metals (Cr, Nb, Mo, Pd).</li>\n</ul>\n<h3 id=\"h2-nested-constraint-hierarchy\">H2 — Nested constraint hierarchy</h3><p>The binding constraint for 3d/4d metals is the XC functional; for 5d metals a deeper physical constraint (e.g., scalar relativistic / correlation effects) may supersede the XC functional.</p>\n<ul>\n<li><strong>Prediction 2a:</strong> 3d/4d subset clusters significantly by functional (exact permutation p &lt; 0.05).</li>\n<li><strong>Prediction 2b:</strong> The 5d metals with a PBE baseline (Pt, W) do not cluster by functional (p &gt; 0.20) and PBE-to-r2SCAN error vectors maintain high cosine similarity (&gt; 0.8). <em>Au is included in the target set but uses a PW91-GGA baseline, so it is excluded from the PBE-vs-r2SCAN functional-clustering test; Pt and W are the only PBE-baseline 5d metals.</em></li>\n<li><strong>Kill condition:</strong> The available PBE-baseline 5d metals cluster strongly by functional (p &lt; 0.05) while the 3d/4d metals do not.</li>\n<li><strong>Scope adjustment:</strong> Because Au lacks a stable PBE tensor, the original “5d noble metals (Au, Pt)” functional-clustering subset is revised to “PBE-baseline 5d metals (Pt, W)”; Au is retained in the target set with a PW91 baseline.</li>\n</ul>\n<h3 id=\"h3-rotation-link-to-layer-3\">H3 — Rotation link to Layer 3</h3><p>The empirical XC bias vector (T_r2SCAN − T_PBE) aligns directionally with pseudopotential-based DFT error vectors from Layer 3.</p>\n<ul>\n<li><strong>Kill condition:</strong> Cosine similarity between Layer 2 XC bias vector and Layer 3 PBE DFT error vector &lt; 0.5 for the majority of elements.</li>\n<li><strong>Status:</strong> Layer 3 DFT compute is staged (see <code>replication/error-geometry/prereg_r2b_dft_anchor_spec.md</code>). This hypothesis remains pending until the all-electron anchor runs complete.</li>\n</ul>\n<h3 id=\"h4-compute-budget-head-to-head\">H4 — Compute-budget head-to-head</h3><p>For elastic-constant prediction, one MLIP run + the Lupine Correction Operator achieves lower out-of-sample MSE than the mean of a 4-model ensemble, with tighter conformal-calibrated intervals.</p>\n<ul>\n<li><strong>Kill condition:</strong> MSE(Operator) ≥ MSE(ensemble mean) or conformal coverage &lt; 90%.</li>\n<li><strong>Evaluated-subset note:</strong> On the 6-element classical-potential benchmark, the operator won 4/6 head-to-head comparisons (Cu and Pt went to ensemble). Conformal coverage was ≥ 90% on all elements. This is reported as interim evidence, not a definitive test of the full Layer-2 MLIP claim.</li>\n</ul>\n<h2 id=\"analysis-plan\">Analysis plan</h2><ol>\n<li><strong>Geometry:</strong> compute participation ratio (PR) of each ensemble error matrix; expect PR ≈ 1.0–1.3.</li>\n<li><strong>Bias extraction:</strong> first principal component of the centered error matrix = 1D bias vector <code>b</code>.</li>\n<li><strong>Functional shift:</strong> Δf = T_r2SCAN − T_PBE.</li>\n<li><strong>Operator:</strong> <code>corrected = raw − b + Δf</code>.</li>\n<li><strong>Uncertainty:</strong> split-conformal prediction on leave-one-out residuals; report 90% coverage and interval width.</li>\n<li><strong>Significance:</strong> exact permutation tests for functional clustering; report p-values and effect sizes.</li>\n</ol>\n<h2 id=\"software-artifacts\">Software artifacts</h2><ul>\n<li><code>lammps-operator/lupine_operator.py</code> — Projection Law operator.</li>\n<li><code>lammps-operator/lammps_harness.py</code> — deterministic 0K elastic-constant harness.</li>\n<li><code>lammps-operator/run_benchmark.py</code> — head-to-head benchmark orchestrator.</li>\n<li><code>lammps-operator/curate_targets.py</code> — target curation script.</li>\n<li><code>data/targets_0K.json</code> — ground-truth target values (14 elements, <code>lupine.targets_0K.v2</code>).</li>\n<li><code>data/pbe_targets_dejong2015.json</code> — raw PBE reference values (14 elements).</li>\n<li><code>data/curate_targets_0K.py</code> — script that applies the scalar bulk-modulus shift and stability gate.</li>\n</ul>\n<h2 id=\"scientific-integrity-policy\">Scientific-integrity policy</h2><ul>\n<li>No synthetic data in published claims.</li>\n<li>Every <code>BenchmarkEntry.predicted</code> must carry a <code>LammpsRun</code> provenance record.</li>\n<li>Every theorem about computed values must use <code>native_decide</code> or <code>by decide</code> in Lean; no <code>rfl</code> on floats.</li>\n<li>Build failures in <code>#guard</code> statements are treated as scientific discrepancies.</li>\n</ul>\n"}