{"id":"z3-adsorption-delta-campaign-round4","title":"Z3 Δ-Learned Adsorption Accuracy — Round-4 Results","subtitle":"Refuted — 128/128 baseline measurements; every validation-selected correction worsened the holdout (2.27–5.91 eV vs the 0.1 eV gate).","category":"validation","tags":["z3","adsorption","delta-learning","catbench","round-4","featured"],"source":"articles/docs/validation/z3-adsorption-delta-campaign-round4-results.md","lang":"en","words":758,"readMinutes":3,"toc":[{"depth":2,"text":"Preregistered question","id":"preregistered-question"},{"depth":2,"text":"Panel and protocol","id":"panel-and-protocol"},{"depth":2,"text":"Baseline result: systematic underbinding, panel-wide","id":"baseline-result-systematic-underbinding-panel-wide"},{"depth":2,"text":"Δ-correction: fitted, selected, scored — and refuted","id":"correction-fitted-selected-scored-and-refuted"},{"depth":2,"text":"What stands after refutation","id":"what-stands-after-refutation"},{"depth":2,"text":"Receipts","id":"receipts"}],"html":"<h1 id=\"z3-learned-adsorption-accuracy-round-4-campaign-results\">Z3 Δ-Learned Adsorption Accuracy — Round-4 Campaign Results</h1><p><strong>Status:</strong> completed campaign · verdict <strong>refuted</strong> (as preregistered)\n<strong>Campaign:</strong> <code>discovery.round-4.z3-adsorption.v1</code> · executed 2026-07-19 (run <code>z3-20260719</code>)\n<strong>Gate:</strong> corrected adsorption-energy MAE ≤ 0.1 eV against published DFT references on a 20-candidate holdout\n<strong>Verdict:</strong> all four available models FAIL — corrected holdout MAE 2.27–5.91 eV; every validation-selected correction made the holdout <em>worse</em> than the raw baseline</p>\n<h2 id=\"preregistered-question\">Preregistered question</h2><p>Can a Δ-learned hybrid stack — a foundation MLIP plus a small correction model fitted on six chemistry-family training systems and selected on six validation systems — reach the ≤ 0.1 eV adsorption-energy accuracy catalyst screening requires, on 20 held-out adsorbate–surface pairs it never saw during fitting, selection, or tuning?</p>\n<h2 id=\"panel-and-protocol\">Panel and protocol</h2><ul>\n<li><strong>Reference panel:</strong> <code>data/candidates/z3_catbench_bm_adsorption.lock.json</code> (SHA-256 <code>b434de00…</code>) — 32 rows from the CatBench BM_dataset adsorption benchmark (Zenodo DOI <a href=\"https://doi.org/10.5281/zenodo.17157086\">10.5281/zenodo.17157086</a>, CC BY 4.0), structures and energies from the GAME-Net study (DOI <a href=\"https://doi.org/10.1038/s43588-023-00437-y\">10.1038/s43588-023-00437-y</a>; VASP 5.4.4 PBE+D2, 450 eV cutoff, PAW, 1e-5 eV SCF, 0.03 eV/Å forces). fcc(111) facets for Ag/Au/Cu/Ni/Pt, hcp(0001) for Ru. Three application families: biomass, plastics, polyurethanes.</li>\n<li><strong>Basis honesty:</strong> references are published <strong>DFT</strong>, not experiment; nothing here is error-against-experiment. Per-row <code>uncertainty_ev</code> is a labeled 3×SCF-threshold proxy, not a statistical interval.</li>\n<li><strong>Frozen split:</strong> deterministic, family-stratified by SHA-256 ordering — 6 <code>delta_train</code> / 6 <code>delta_validation</code> / 20 <code>confirmatory_test</code> (<code>z3_catbench_bm_delta_splits.lock.json</code>). Fit exclusion enforced in code: fit reads train only, selection reads validation only, scoring reads test only.</li>\n<li><strong>Execution:</strong> 4 models × 32 candidates = 128 single-candidate Cloud Run cells (isolated jobs, defect-fixed images <code>z3-adsorption-fixed-20260719</code>, service account <code>atlas-distill-runner</code>). <strong>128/128 completed, zero failures</strong>, each artifact captured and identity-validated (model, row, candidate, finite energy) by <code>gcp/z3-campaign/run_measurement.py</code>.</li>\n</ul>\n<h2 id=\"baseline-result-systematic-underbinding-panel-wide\">Baseline result: systematic underbinding, panel-wide</h2><p>Signed error = model − reference (positive = underbound). Holdout = the 20 confirmatory candidates.</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Model</th>\n<th>Full-panel MAE</th>\n<th>Full-panel bias</th>\n<th>Holdout MAE</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Model\">chgnet</td>\n<td data-label=\"Full-panel MAE\">1.67 eV</td>\n<td data-label=\"Full-panel bias\">+1.13 eV</td>\n<td data-label=\"Holdout MAE\"><strong>0.69 eV</strong></td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mp-medium</td>\n<td data-label=\"Full-panel MAE\">3.43 eV</td>\n<td data-label=\"Full-panel bias\">+3.42 eV</td>\n<td data-label=\"Holdout MAE\">2.11 eV</td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mp-small</td>\n<td data-label=\"Full-panel MAE\">4.29 eV</td>\n<td data-label=\"Full-panel bias\">+4.28 eV</td>\n<td data-label=\"Holdout MAE\">3.24 eV</td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mpa-0-medium</td>\n<td data-label=\"Full-panel MAE\">5.31 eV</td>\n<td data-label=\"Full-panel bias\">+5.29 eV</td>\n<td data-label=\"Holdout MAE\">4.27 eV</td>\n</tr>\n</tbody></table></div><p>Errors are not symmetric scatter: the MACE family underbinds essentially every system, with per-candidate errors from −1.1 eV to <strong>+25.6 eV</strong>, growing with adsorbate size and varying by family (plastics largest, polyurethanes smallest). The physical reading is dispersion blindness: the references stabilize large adsorbates partly through D2 dispersion, physics the foundation models&#39; bulk-crystal training distribution does not contain. First datapoint anatomy (mace-mp-small × biomass_ni_mol1): gas molecule +1.39 eV, clean Ni(111) slab −12.45 eV, complex −1.24 eV — the slab/complex terms nearly cancel and the residual <strong>+9.8 eV sits entirely in the interface bond</strong>.</p>\n<h2 id=\"correction-fitted-selected-scored-and-refuted\">Δ-correction: fitted, selected, scored — and refuted</h2><p>A fixed correction menu was fitted per model on <code>delta_train</code> only and selected on <code>delta_validation</code> only (<code>tools/build_z3_delta_correction.py</code>; report <code>data/candidates/z3/delta-correction-report.json</code> + <code>.sha256</code>):</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Model</th>\n<th>Selected form (by validation MAE)</th>\n<th>Validation MAE</th>\n<th>Baseline holdout MAE</th>\n<th><strong>Corrected holdout MAE</strong></th>\n<th>Gate</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Model\">chgnet</td>\n<td data-label=\"Selected form (by validation MAE)\">C: family size-linear</td>\n<td data-label=\"Validation MAE\">1.12 eV</td>\n<td data-label=\"Baseline holdout MAE\">0.69 eV</td>\n<td data-label=\"Corrected holdout MAE\"><strong>2.27 eV</strong></td>\n<td data-label=\"Gate\">✗</td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mp-medium</td>\n<td data-label=\"Selected form (by validation MAE)\">A: global constant</td>\n<td data-label=\"Validation MAE\">3.26 eV</td>\n<td data-label=\"Baseline holdout MAE\">2.11 eV</td>\n<td data-label=\"Corrected holdout MAE\"><strong>5.01 eV</strong></td>\n<td data-label=\"Gate\">✗</td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mp-small</td>\n<td data-label=\"Selected form (by validation MAE)\">A: global constant</td>\n<td data-label=\"Validation MAE\">3.98 eV</td>\n<td data-label=\"Baseline holdout MAE\">3.24 eV</td>\n<td data-label=\"Corrected holdout MAE\"><strong>5.00 eV</strong></td>\n<td data-label=\"Gate\">✗</td>\n</tr>\n<tr>\n<td data-label=\"Model\">mace-mpa-0-medium</td>\n<td data-label=\"Selected form (by validation MAE)\">B: family constant</td>\n<td data-label=\"Validation MAE\">4.03 eV</td>\n<td data-label=\"Baseline holdout MAE\">4.27 eV</td>\n<td data-label=\"Corrected holdout MAE\"><strong>5.91 eV</strong></td>\n<td data-label=\"Gate\">✗</td>\n</tr>\n</tbody></table></div><p><strong>Every corrected holdout is worse than its own raw baseline.</strong> A global shift overshoots the small-error families into large negative errors; family/size-linear fits trained on two giant plastics systems extrapolate badly across the size distribution. The conclusion is structural, not anecdotal: the baseline error is <em>not a uniform bias</em> — it is a structured, family- and size-dependent field spanning −1 to +26 eV — and a six-point fit budget cannot estimate a generalizable correction over it. The Z3 Δ-learning hypothesis <strong>as preregistered</strong> is refuted on the holdout it froze.</p>\n<h2 id=\"what-stands-after-refutation\">What stands after refutation</h2><ul>\n<li><strong>chgnet raw</strong> (0.69 eV holdout MAE) is the best bare foundation-model number on this panel — still 6.9× the screening gate. No current available uMLIP is catalyst-screening accurate on biomass/plastics-scale adsorbates.</li>\n<li>The underbinding law is now measured on <strong>three independent observables</strong> in this program: Z1 migration barriers (3.4–6× under-predicted), Z3 adsorption interfaces (up to +26 eV underbound), Round-4 elastic correction (confirmatory fail). One systematic direction — underbinding at transition states and at interfaces — across three preregistered campaigns.</li>\n<li>A viable corrected future run needs a materially larger fit budget (more than six train systems) or physics features (contact-atom counts, dispersion proxies) — a new preregistration, not a retune of this one.</li>\n</ul>\n<h2 id=\"receipts\">Receipts</h2><ul>\n<li>Manifest: <code>campaigns/v1/z3.campaign-manifest.v1.json</code> (content hash <code>sha256:49f5f20e…</code>); claim <code>registry/claims/discovery.z3.adsorption-accuracy.v1.json</code> (unsupported; baseline bundle PR #39).</li>\n<li>Panel/splits/fixtures: <code>data/candidates/z3_catbench_bm_*</code> (all <code>.sha256</code> sidecars); 32 candidate fixtures <code>gs://shed-489901-atlas-inputs/z3-campaign/catbench-bm-v1/</code>.</li>\n<li>Raw artifacts: <code>gs://shed-489901-atlas-outputs/z3-campaign/raw/z3-20260719/&lt;model&gt;/&lt;candidate&gt;/adsorption_energy/cell_result.json</code>.</li>\n<li>Analysis: <code>tools/build_z3_delta_correction.py</code>; report <code>data/candidates/z3/delta-correction-report.json</code> (SHA-256 sidecar); 7 tests covering form recovery, fallback guards, no-leakage, fail-closed hashing, gate arithmetic.</li>\n<li>Runner defects found and fixed before execution (PR #28): failed-checkpoint reuse; raw-energy loss on contribution overflow.</li>\n</ul>\n"}