{"id":"environment-error-field-paper","title":"The Environment Error Field — Manuscript","subtitle":"uMLIP property errors as projections of a smooth coordination field: blind r=0.906, run-time correction, machine-verified claims incl. the nulls.","category":"validation","tags":["mlip","y-matrix","environment-field","isotonic","lean","live-lab"],"source":"articles/paper/environment-error-field-2026-07-02.md","lang":"en","words":2647,"readMinutes":12,"toc":[{"depth":2,"text":"Abstract","id":"abstract"},{"depth":2,"text":"1. Introduction","id":"1-introduction"},{"depth":2,"text":"2. Results","id":"2-results"},{"depth":3,"text":"2.1 The error landscape (Fig. 1)","id":"2-1-the-error-landscape-fig-1"},{"depth":3,"text":"2.2 Rankings survive where magnitudes fail (Fig. 2)","id":"2-2-rankings-survive-where-magnitudes-fail-fig-2"},{"depth":3,"text":"2.3 The error is approximately log-affine; the exponent orders by family (Fig. 3)","id":"2-3-the-error-is-approximately-log-affine-the-exponent-orders-by-family-fig-3"},{"depth":3,"text":"2.4 The environment error field and its blind test (Fig. 4)","id":"2-4-the-environment-error-field-and-its-blind-test-fig-4"},{"depth":3,"text":"2.5 From field to run time: an energy-level correction (Fig. 5)","id":"2-5-from-field-to-run-time-an-energy-level-correction-fig-5"},{"depth":3,"text":"2.6 Provable boundaries (kernel-checked)","id":"2-6-provable-boundaries-kernel-checked"},{"depth":2,"text":"3. Discussion","id":"3-discussion"},{"depth":2,"text":"4. Methods","id":"4-methods"},{"depth":3,"text":"4.1 The property matrix","id":"4-1-the-property-matrix"},{"depth":3,"text":"4.2 References","id":"4-2-references"},{"depth":3,"text":"4.3 Statistics","id":"4-3-statistics"},{"depth":3,"text":"4.4 Machine verification","id":"4-4-machine-verification"},{"depth":3,"text":"4.5 Literature verification","id":"4-5-literature-verification"},{"depth":2,"text":"Data, code, and proof availability","id":"data-code-and-proof-availability"}],"html":"<h1 id=\"a-smooth-environment-resolved-error-field-underlies-the-systematic-property-errors-of-universal-machine-learned-interatomic-potentials\">A smooth, environment-resolved error field underlies the systematic property errors of universal machine-learned interatomic potentials</h1><blockquote>\n<p><strong>Draft v0.2 — 2026-07-02.</strong> Replaces v0.1 after statistical hardening,\nfigure generation, and bibliography verification. §2.5 awaits the run-level\ncorrection experiment (slots marked <code>[PENDING-RUN]</code>); all other numbers are\nfinal. Citations <code>[key]</code> resolve in <code>paper/references-envfield.md</code> (77\nverified entries). Figures in <code>paper/figures/envfield/</code> with SHA-256 input\nmanifest. Statistical methods and weakened-claim registers follow\n<code>data/y_matrix_runs/analysis/STATS.md</code> exactly.</p>\n</blockquote>\n<h2 id=\"abstract\">Abstract</h2><p>Universal machine-learned interatomic potentials (uMLIPs) err systematically\non the properties that dominate materials practice — surfaces, vacancies,\nplanar faults — while matching references closely on bulk observables\n[deng2024softening, focassio2025surfaces, chipsff2025]. We show that for\nface-centered-cubic metals these errors are consistent with projections of a\nsingle object: a smooth error field over local atomic environments, with\ncoordination deficit as leading coordinate. Measured per (model, material)\nfrom three standard observables (γ₁₀₀, γ₁₁₁, and the vacancy formation\nenergy), the field predicts a fourth, never-fitted observable that probes an\nunfitted coordination: across 36 (model, material) cells the blind γ₁₁₀\nprediction attains r = 0.906, exceeding all 10,000 within-model material\npermutations (p = 10⁻⁴ against a null whose mean is 0.44, not zero;\nmaterial-clustered 95% CI [0.82, 0.96]), with zero adjustable parameters.\nThe field organizes previously disconnected observations: near-universal\npreservation of cross-material rankings amid large magnitude errors; an\napproximately log-affine error form whose exponent orders by property family\nwithin every model while training lineage moves the prefactor toward unity;\nand the failure of scalar corrections at every grain. Because the field is a\nfunction of environments, its inverse is an additive energy with analytic\nforces, deployed here beside a live CHGNet calculator: it recovers the\nfitted observables exactly through full relaxations, improves the blind\nfacet at run time (γ₁₁₀ error 9.7 → 1.5 % for Ni, 28.0 → 13.7 % for Cu),\nleaves the bulk structurally untouched, and runs stable molecular dynamics\nat 15.6 % overhead — while leaving near-equilibrium force errors unchanged,\na null result that cleanly scopes the first-shell field as an energy-level\ncorrection and identifies the continuous-coordinate extension required for\nforces. Where rankings invert, we prove — machine-checked — that no\nmonotone correction exists.\nAll quantitative claims, including negative ones, are sealed as\nproof-assistant-verified theorems over provenance-hashed data.</p>\n<h2 id=\"1-introduction\">1. Introduction</h2><p>Foundation uMLIPs [batatia2022mace, deng2023chgnet, chen2022m3gnet,\nmace-mp-0] bring near-DFT accuracy to million-atom simulation, and community\nbenchmarks document rapid progress on bulk energetics and stability\n[matbench-discovery, matcalc-2025]. The same benchmarks, and dedicated\nstudies, document persistent failures away from equilibrium: surface\nenergies under-predicted by every MPtrj-trained model\n[focassio2025surfaces], elastic tensors and equations of state markedly\nharder than relaxed geometries [chipsff2025], and a pervasive &quot;softening&quot; of\nthe potential-energy surface traced to near-equilibrium bias in training\ndata [deng2024softening, omat24]. Two questions remained open in the\nliterature we could verify (§4.5): whether cross-property error structure\nexists for foundation models — the closest study concerns classical\npotentials of a single element [ni47metrics] — and whether a correction\nfitted on one property class transfers to another; no\ntransferability-of-correction study has been published to our knowledge.</p>\n<p>Here we answer both for fcc metals, and the answers compose. Cross-property\nerror structure exists, but not where prior tools looked: it is absent in\nlinear property-space statistics (pre-registered participation-ratio and\nmode-alignment tests fall inside coupling-aware nulls) and present as a\n<em>smooth field over local atomic environments</em>. The field is measurable from\nthree observables, transfers blind to a fourth (§2.4), converts into a\nrun-time force-bearing correction (§2.5), and carries provable applicability\nboundaries (§2.6). Throughout, claims are sealed as machine-checked theorems\n(§4.4); we report one instance where the proof kernel rejected a claim that\nhad survived our statistical filter, and the corrected count.</p>\n<h2 id=\"2-results\">2. Results</h2><h3 id=\"2-1-the-error-landscape-fig-1\">2.1 The error landscape (Fig. 1)</h3><p>We evaluated four uMLIPs — CHGNet 0.4.2 [deng2023chgnet], MACE-MP-0 small\nand medium [mace-mp-0], and MACE-MPA-0 (OMat24 lineage) [omat24] — on a\n21-material × up-to-9-property matrix (§4.1) against 228\nprovenance-annotated published references (§4.2). Bulk observables are\naccurate (median |relative error|: lattice constants &lt; 0.5 %, formation\nenthalpies ≈ 3 %); defect-family observables err 15–60× worse per model\n(bootstrap CIs exclude parity for all four; Fig. 1b). This split — the\ndefect/bulk asymmetry of [deng2024softening] quantified on a matched,\nreference-bound matrix — motivates everything that follows.</p>\n<h3 id=\"2-2-rankings-survive-where-magnitudes-fail-fig-2\">2.2 Rankings survive where magnitudes fail (Fig. 2)</h3><p>Across materials, predicted rankings track reference rankings closely for\nsurfaces (Spearman ρ = 0.88–1.00 per model; MACE-MPA-0&#39;s γ₁₁₁ ranking\nreproduces the reference permutation exactly), vacancies (0.84–0.93), and\nB₀ (0.82–0.85), while the corresponding magnitude errors reach tens of\npercent. All 22 reference-ordered facet hierarchies are reproduced by all\nmodels. The single fracture is diagnostic: stacking-fault rankings collapse\nfor the three MPtrj-trained models (ρ = 0.11–0.46) and survive in the\nOMat-lineage model (0.93) — we return to why in §2.4. Ordinal faithfulness\nis the invertibility condition for any monotone error model; its selective\nfailure marks where no such model can apply.</p>\n<h3 id=\"2-3-the-error-is-approximately-log-affine-the-exponent-orders-by-family-fig-3\">2.3 The error is approximately log-affine; the exponent orders by family (Fig. 3)</h3><p>Regressing log-prediction on log-reference yields R² = 0.93–0.98 (surfaces,\nB₀): the error acts as pred ≈ c·T^α within a property family. Two\nregularities follow, stated in the registers our statistics support (§4.3).\n<em>Family ordering:</em> in all four models the fitted surface exponent exceeds\nthe vacancy and B₀ exponents (8/8 point-orderings, a deterministic property\nof this dataset); as statistical claims, 5/8 paired-bootstrap differences\nexclude zero nominally and 1/8 after Holm correction. <em>Cross-model\ncompatibility:</em> the four surface exponents (1.065–1.138) spread less\nbetween models than single-model uncertainty (variance ratio 0.39) —\nconsistent with, though not demonstrating, a family-owned exponent.\nTraining lineage acts on the prefactor: surfaces c = 0.66 (CHGNet) → 0.98\n(MPA-0), and pooled warp magnitude orders strictly by lineage\n(0.051 &lt; 0.099 &lt; 0.120 &lt; 0.388). A two-parameter log-affine correction\nmatches an 8-knot isotonic correction out-of-sample (paired-difference CIs\ninclude zero) with six fewer parameters; against raw predictions the\ncorrection is decisive for CHGNet (27.96 % → 10.04 %, clustered\np = 4×10⁻⁴) and directional for MACE-small (12.05 % → 7.45 %, p = 0.061).\nScalar (α = 1) corrections fail at every grain we tested — per-model,\nper-material, within-family — consistent with α ≠ 1.</p>\n<h3 id=\"2-4-the-environment-error-field-and-its-blind-test-fig-4\">2.4 The environment error field and its blind test (Fig. 4)</h3><p>The regularities of §2.2–2.3 follow if the model&#39;s energy error is a smooth\nfunction of local coordination, accumulated per atom:</p>\n<p>  E_model(config) − E_ref(config) ≈ Σᵢ Δε(cᵢ),   Δε(12) ≡ 0 (fcc bulk).  (1)</p>\n<p>Each property then samples the field at its characteristic coordinations —\nfcc(100) top-layer atoms at c = 8, (111) at 9, vacancy first neighbors at\n11, (110) at 7 and 11 — so per (model, material) three observables measure\nthree field values:</p>\n<p>  Δε(8) = δγ₁₀₀·A₁₀₀,  Δε(9) = δγ₁₁₁·A₁₁₁,  Δε(11) = δE_vac/12,  (2)</p>\n<p>with δ the signed error and A the area per surface atom. The field\nhypothesis is then falsifiable with no free parameters: γ₁₁₀&#39;s error\ninvolves the <em>unfitted</em> coordination 7, predicted by linear continuation of\nthe field below c = 8. Across 36 cells the prediction attains r = 0.906\n(Fig. 4b) — exceeding all 10,000 within-model material permutations\n(p = 10⁻⁴; the honest null has mean r = 0.44 because pooling across models\nshares error scales, and we report against it, not against zero) — with\nmaterial-clustered 95% CI [0.82, 0.96] and no single material carrying the\nresult (leave-one-material-out r = 0.857–0.944). Per model, the prediction\nis individually significant for CHGNet (r = 0.86), MACE-small (0.90), and\nMPA-0 (0.96), but not MACE-medium (0.47, p = 0.10, n = 9): the field claim\nis carried by three of the four models. The median residual improves on\npredict-zero marginally under clustering (0.066 vs 0.104 J/m², p = 0.036);\n26 of 36 cells improve strictly at 10⁻⁴ J/m² integer precision (one\nadditional cell&#39;s improvement vanishes at that precision — a margin caught\nby the proof kernel, §4.4).</p>\n<p>The field explains the fracture of §2.2: an intrinsic stacking fault alters\nno first-neighbor counts, so a first-shell field is blind to it — SFE\nerrors are ungoverned residue, and MPtrj models, whose training\ndistribution samples faulted stackings sparsely, scramble SFE rankings\nwhile everything first-shell-visible stays ordered.</p>\n<h3 id=\"2-5-from-field-to-run-time-an-energy-level-correction-fig-5\">2.5 From field to run time: an energy-level correction (Fig. 5)</h3><p>Equation (1) inverts into an additive correction energy\nE_corr = −Σᵢ P(cᵢ), with P the cubic through the three measured knots and\nP(12) = 0, cᵢ a smooth-cutoff coordination, and analytic forces via the\nchain rule (verified against numerical differentiation to 10⁻⁶ eV/Å on\nrattled slabs), deployed beside the live CHGNet calculator. Three\nvalidation levels, as measured:</p>\n<p><em>Statics.</em> The corrected calculator, run through the full property\npipeline (relaxations included), recovers the three fitted observables\nexactly — a closure test showing the bond-counting field survives real\nrelaxation — and improves the <strong>blind</strong> facet at run time: γ₁₁₀ error\n9.7 % → 1.5 % (Ni) and 28.0 % → 13.7 % (Cu). The bulk is structurally\nuntouched: coordination remains 12 across the EOS window, so E_corr\nvanishes identically there (measured a₀ shift 0.0000 Å).</p>\n<p><em>Forces (null result).</em> Force RMSE against a stronger-model proxy\n(MACE-MPA-0) on twenty rattled Ni(110) slabs is unchanged (0.2594 →\n0.2595 eV/Å overall; 0.1926 → 0.1932 on surface atoms). The mechanism is\ninstructive: the first-shell field corrects energy <em>differences between\ncoordination environments</em>, and its force footprint is confined to the\ncutoff switching shell, which 0.08 Å thermal displacements do not cross.\nNear-equilibrium force error lives in the PES curvature — the softening\nexponent of §2.3 — not in the coordination step. Version 1 is therefore\nan <strong>energy-level</strong> correction (property values, and by extension\nenergy differences between coordination states); correcting forces\nrequires a field over a continuous environment coordinate, which we\nidentify as the necessary v2.</p>\n<p><em>Dynamics.</em> 1,000 steps of 300 K Langevin NVT on the corrected Ni(110)\nslab run stably (the ≈5 eV total-energy rise matches 3/2·N·k_BT\nequilibration from a cold start); the unoptimized correction overlay adds\n15.6 % wall time to the CHGNet step, with an obvious path to negligible\ncost via a compiled pairwise implementation (the term is EAM-embedding\nshaped and LAMMPS-overlay compatible [lammps-docs]).</p>\n<h3 id=\"2-6-provable-boundaries-kernel-checked\">2.6 Provable boundaries (kernel-checked)</h3><p>Correction has jurisdiction only where order survives. Where it does not,\nwe prove impossibility rather than report failure: MACE-MP-small orders\nSFE(Ni) ≤ SFE(Al) while the references order the reverse; a quantified\nmonotonicity lemma then shows no monotone correction maps both predictions\nto their references (machine-checked with concrete witnesses). Two further\nboundaries: separately-fitted classical-potential families share no field —\nleave-element-out calibration of 623 ledger EAM elastic records <em>degrades</em>\nthem — and already-converged cells refuse correction (MPA-0&#39;s surfaces,\nwhere raw error sits at the anchor-noise floor). Classical potentials\nretain their own value proposition: a consistent EAM family beats CHGNet on\n20/24 matched surface cells at ~6× less compute without a GPU\n(kernel-checked cell facts), a deployment-relevant baseline for\ncorrection economics.</p>\n<h2 id=\"3-discussion\">3. Discussion</h2><p><strong>Relation to prior corrections.</strong> Δ-machine-learning [ramakrishnan2015]\nand its descendants — from coupled-cluster corrections [bogojeski2020,\nnandi2021, zheng2021aiqm1, oneill2025] to uMLIP fine-tuning [radova2025,\nkaur2025, steels2025, huang2025crossfunctional] — <em>learn</em> corrections from\nper-system reference data and yield new weights with no statement of\napplicability. The present correction is <em>measured</em>, not learned: three\nanchor observables fix a closed-form field; transfer within the family is\nthe tested content of the method (γ₁₀₀ → γ₁₁₀ blind, §2.4); and\napplicability is gated by proof (§2.6). The one-data-point rescaling of\n[deng2024softening] is the field&#39;s zeroth mode — a constant Δε — and our\nmeasurements show why it saturates: the error is environment-resolved\n(15–60× defect/bulk asymmetry) and its log-affine exponent differs from\nunity. Conversely, wherever abundant system-specific DFT is affordable,\nfine-tuning strictly dominates, including in the order-scrambled cells our\ngates refuse. Output-space calibration — the isotonic lineage\n[ayer1955, barlow1972, zadrozny2002, guo2017] and materials-UQ\nrecalibration [tran2020uq, pernot2022] — corrects numbers after the run;\nEq. (1) corrects forces during it.</p>\n<p><strong>Mechanism (hypothesis).</strong> A smooth regressor trained on a\nconfiguration distribution dominated by near-equilibrium, high-coordination\nenvironments will interpolate confidently there and extrapolate with\ncorrelated bias into under-sampled coordination regimes; the bias, shared\nacross materials because the descriptor space is shared, appears as a\nsmooth Δε(c). This is testable independently of our data: the coordination\nhistogram of MPtrj vs OMat24 [mptrj, omat24] should predict the field&#39;s\nmagnitude, and the sloppy-model geometry of the fitted models\n[transtrum2010, transtrum2011, kurniawan2022] should exhibit the\ncorresponding stiff direction. We propose both as follow-up.</p>\n<p><strong>Relation to the error-geometry program.</strong> Pre-registered linear analyses\n(participation ratios, mode cosines against coupling-preserving nulls)\nfound no cross-property structure — apparent cosine alignments of 0.96 fell\ninside nulls reaching 0.98, a false positive avoided only by the null\ndesign. The structure is nonlinear: one smooth function per (model,\nmaterial, family). Low-dimensionality of model error [transtrum2010]\nsurvives, in curved form, one level below the observables.</p>\n<p><strong>Limitations.</strong> Coverage: fcc first-shell coordination only; bcc and\nhcp surfaces, alloys, and finite-temperature observables are untested;\nthe field needs at least a second-shell coordinate to govern planar\nfaults. Statistics: n = 9 materials per model for the blind test; one of\nfour models is individually non-significant; exponent family-ordering is\nstatistically resolved for only a subset after multiplicity correction.\nPhysics: relaxation contributions are folded into effective knots\n(unrelaxed bond counting); vacancy references mix DFT functionals; the\nforce validation uses a stronger model as proxy, not DFT, and returned a\nnull — the v1 field does not correct near-equilibrium forces (§2.5), so\nclaims of run-time benefit are limited to energetics until a\ncontinuous-coordinate field is built and tested.\nScope of verification: the proof kernel certifies data-analysis\narithmetic and stated inequalities over the measured dataset — it does\nnot certify physics.</p>\n<h2 id=\"4-methods\">4. Methods</h2><h3 id=\"4-1-the-property-matrix\">4.1 The property matrix</h3><p>21 materials (Ag, Al, Au, Ca, Cr, Cu, Fe, Mo, Nb, Ni, Pd, Pt, Sr, Ta, V, W;\nSi; B2-NiAl; L1₂-Ni₃Al; MgO; NaCl) × 4 models × up to 9 properties (a₀,\nB₀, B₀′, E_vac, γ₁₀₀, γ₁₁₀, γ₁₁₁, γ_SFE, ΔH_f). Statics: Birch–Murnaghan\nEOS (±6 % volume, 11 points, recentring); fixed-cell FIRE relaxation\n(fmax 0.01 eV/Å); slabs ≥ 8 layers, 12 Å vacuum; displaced-slab intrinsic\nSFE evaluating both Shockley branches; 3×3×3 vacancy supercells. Each\ncell is provenance-hashed and bit-reproducible (≤ 10⁻¹¹). Fe carries a\ndocumented model-deficiency annotation (its anomalous B₀ is\nwindow-independent and unaffected by magnetic initialization, which we\nverified changes neither calculator&#39;s energy by even one ulp; an earlier\nstate-preparation hypothesis of ours is preserved, falsified, in the\nregistration&#39;s amendment log).</p>\n<h3 id=\"4-2-references\">4.2 References</h3><p>228 published values, 8 property families, compiled under a no-fabrication\nrule: every value carries a citation verified at compilation time\n[tran2016, dejong2015, angsten2014, ma-dudarev2019, ...20 primaries in\nreferences file]; unverifiable values are recorded as explicit gaps.\nBinding prefers DFT-PBE, falls back to experiment, and mechanically\nrefuses method-mismatched entries.</p>\n<h3 id=\"4-3-statistics\">4.3 Statistics</h3><p>Pre-registered hypotheses with kill conditions; coupling-aware nulls\n(within-family structure preserved, cross-family alignment permuted;\n1,000 seeded draws); leave-one-out throughout; material-clustered\nbootstrap (10,000 draws) for all headline CIs; within-model material\npermutation (10,000 draws) for the blind-test null; Holm step-down within\npre-specified claim families; ~47 sampling-based tests inventoried, with\nthe primary blind-test claim surviving paper-wide Bonferroni. Weakened\nregisters adopted wherever hardening reduced a claim; the full audit is\nreleased with the data.</p>\n<h3 id=\"4-4-machine-verification\">4.4 Machine verification</h3><p>All quantitative outcomes — positive, negative, and impossible — are\nencoded as decidable Lean 4 theorems over integer-scaled data carrying\nSHA-256 provenance of their sources: ordering claims as inequality chains\nover predictions (the kernel verifies the order, not a summary), the\nisotonic correction as a Lean function whose outputs the kernel computes,\nimpossibility results as quantified lemmas with concrete witnesses. Eight\nmodules, 100+ theorems, zero <code>sorry</code>. During preparation the kernel\nrejected one claimed blind-prediction win whose margin vanished at\ninteger precision (<code>decide</code> refused <code>813 &lt; 813</code>); the count in §2.4 is\nthe corrected one, and the figure pipeline independently reproduced the\nsame boundary case.</p>\n<h3 id=\"4-5-literature-verification\">4.5 Literature verification</h3><p>Two ~100-agent deep-research sweeps with three-vote adversarial\nverification of every extracted claim (25 and 22 claims confirmed 3–0;\nsources and transcripts released). One initially-cited preprint was found\nwithdrawn during bibliography verification and is struck with a dated\ncorrection preserved in the working documents.</p>\n<h2 id=\"data-code-and-proof-availability\">Data, code, and proof availability</h2><p>Evidence payloads, reference compilations with per-value provenance,\nbinding reports, analysis artifacts (seeds included), figure pipeline with\ninput-hash manifest, statistical-hardening script, and all Lean modules\nwith build configuration are available in the project repository; every\ntheorem type-checks under Lean 4.30 with the pinned toolchain.</p>\n"}