{"id":"mlip-elastic-benchmark-protocol","title":"MLIP Elastic Benchmark Protocol","subtitle":"Reusable protocol for the MatPES cubic-metal 16-element elastic-constant benchmark.","category":"validation","tags":["mlip","elasticity","benchmark","protocol","reproducibility"],"source":"articles/mlip-elastic-benchmark/mlip-elastic-benchmark-protocol-2026-06-27.md","lang":"en","words":382,"readMinutes":2,"toc":[{"depth":2,"text":"Claim under test","id":"claim-under-test"},{"depth":2,"text":"Element roster (16)","id":"element-roster-16"},{"depth":2,"text":"Four benchmark arms","id":"four-benchmark-arms"},{"depth":2,"text":"Bias operator (arm B)","id":"bias-operator-arm-b"},{"depth":2,"text":"Cost model","id":"cost-model"},{"depth":2,"text":"Execution steps","id":"execution-steps"},{"depth":2,"text":"Config","id":"config"},{"depth":2,"text":"Kill conditions / caveats","id":"kill-conditions-caveats"}],"html":"<h1 id=\"mlip-elastic-benchmark-protocol-10-cost-reduction-for-mlip-elastic-constant-validation\">MLIP Elastic Benchmark Protocol — 10× Cost Reduction for MLIP Elastic-Constant Validation</h1><blockquote>\n<p><strong>Date:</strong> 2026-06-27<br><strong>Parent task:</strong> <code>t_0266945a</code> (synthesizer)<br><strong>Design source:</strong> <code>lupine-rhizo/docs/plans/mlip-elastic-benchmark-master-2026-06-27.md</code> (binding; this protocol is the runnable translation)</p>\n</blockquote>\n<h2 id=\"claim-under-test\">Claim under test</h2><p>On a 16-element cubic-metal benchmark, a Lupine-corrected single-model 1×1×1 elastic-constant calculation matches the accuracy of a 3×3×3 reference run while costing ~10× fewer CPU-seconds, and outperforms a 3-architecture ensemble while costing ~5× fewer CPU-seconds.</p>\n<h2 id=\"element-roster-16\">Element roster (16)</h2><p><code>Ag Al Au Ca Cr Cu Fe Mo Nb Ni Pd Pt Sr Ta V W</code></p>\n<h2 id=\"four-benchmark-arms\">Four benchmark arms</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Arm</th>\n<th>Label</th>\n<th>Model</th>\n<th>Cell</th>\n<th>Correction</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Arm\">A</td>\n<td data-label=\"Label\"><code>raw-1x1x1</code></td>\n<td data-label=\"Model\">TensorNet/PBE</td>\n<td data-label=\"Cell\">1×1×1</td>\n<td data-label=\"Correction\">None</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">B</td>\n<td data-label=\"Label\"><code>corrected-1x1x1</code></td>\n<td data-label=\"Model\">TensorNet/PBE</td>\n<td data-label=\"Cell\">1×1×1</td>\n<td data-label=\"Correction\">LOO-PCA bias + functional shift</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">C</td>\n<td data-label=\"Label\"><code>ref-3x3x3</code></td>\n<td data-label=\"Model\">TensorNet/PBE</td>\n<td data-label=\"Cell\">3×3×3</td>\n<td data-label=\"Correction\">None</td>\n</tr>\n<tr>\n<td data-label=\"Arm\">D</td>\n<td data-label=\"Label\"><code>ensemble-1x1x1</code></td>\n<td data-label=\"Model\">M3GNet/PBE + CHGNet/PBE + TensorNet/PBE</td>\n<td data-label=\"Cell\">1×1×1</td>\n<td data-label=\"Correction\">None (mean)</td>\n</tr>\n</tbody></table></div><ul>\n<li>QET is deduplicated to TensorNet (byte-identical checkpoint).</li>\n<li>Headline targets are <code>TPBE_0K</code> from <code>targets_0K.json</code>.</li>\n</ul>\n<h2 id=\"bias-operator-arm-b\">Bias operator (arm B)</h2><ol>\n<li>For each element, compute the TensorNet/PBE 1×1×1 error vector: <code>raw − TPBE_0K</code>.</li>\n<li>Leave-one-out: fit the first principal component of the centered error matrix on the other 15 elements; apply that bias to the held-out element.</li>\n<li>Add functional shift <code>Tr2SCAN_0K − TPBE_0K</code>.</li>\n<li>Report mean and median MAE over the 16 LOO-corrected predictions.</li>\n</ol>\n<p>Implementation: <code>lupine/python/lupine/operator.py:correct()</code> and <code>leave_one_out_calibration()</code>.</p>\n<h2 id=\"cost-model\">Cost model</h2><p><code>core_hours = runtime_seconds × n_cores / 3600</code>, with <code>n_cores = 1</code> for per-case CPU-equivalent core-hours. All headline costs are cache-warm. The 1×1×1 matrix is re-run after model download so HuggingFace cache misses do not inflate costs.</p>\n<h2 id=\"execution-steps\">Execution steps</h2><ol>\n<li>Warm model cache: run a 2-element smoke test (Ca, Cu) for all three architectures.</li>\n<li>Run the 16-element 1×1×1 matrix for M3GNet, CHGNet, TensorNet, PBE + r2SCAN (96 cases) using <code>lupine/data/run_mlip_elastic_benchmark_1x1x1_matrix.py</code>.</li>\n<li>Aggregate with <code>lupine-mlip-benchmark/scripts/aggregate.py</code>, which combines the new 1×1×1 results, the existing 3×3×3 16-element grid, and <code>targets_0K.json</code>.</li>\n<li>Verify schema and plausibility with <code>lupine-mlip-benchmark/scripts/verify.py</code>.</li>\n<li>Fill placeholders in the preprint, dashboard, and funder brief using the resulting <code>mlip_elastic_benchmark_results.json</code>.</li>\n</ol>\n<h2 id=\"config\">Config</h2><p>The machine-readable case matrix lives in <code>lupine-mlip-benchmark/config/mlip_elastic_benchmark.yaml</code> (192 cases: 16 elements × 3 architectures × 2 functionals × 2 supercells).</p>\n<h2 id=\"kill-conditions-caveats\">Kill conditions / caveats</h2><ul>\n<li>If <code>MAE(B) &gt; MAE(C) + 1.0 GPa</code>, the corrected small cell does not match the supercell reference.</li>\n<li>If <code>MAE(B) ≥ MAE(D)</code>, the corrected single model does not beat the ensemble; pivot headline to the supercell-independence saving only.</li>\n<li>Report median MAE alongside mean because Cr is a pathological outlier.</li>\n<li>r2SCAN targets are scalar bulk-modulus shifted; Au uses a PW91-GGA fallback; QET≡TensorNet; costs are cache-warm and single-seed.</li>\n</ul>\n"}