{"id":"mrh-monotone-reparameterization","title":"The Monotone Reparameterization Hypothesis","subtitle":"Order right, size wrong: the family-owned warp of foundation-model error, its confirmed predictions, and its registered kill conditions.","category":"conjectures","tags":["mrh","error-geometry","mlip","ordinal"],"source":"articles/docs/mrh-monotone-reparameterization.md","lang":"en","words":1150,"readMinutes":5,"toc":[{"depth":2,"text":"The claim, in one sentence","id":"the-claim-in-one-sentence"},{"depth":2,"text":"Formal statement","id":"formal-statement"},{"depth":2,"text":"Evidence (all measured on the 21-material × 4-model Y-matrix, this repo)","id":"evidence-all-measured-on-the-21-material-4-model-y-matrix-this-repo"},{"depth":2,"text":"Relation to the hyper-ribbon program","id":"relation-to-the-hyper-ribbon-program"},{"depth":2,"text":"Relation to published work (honestly bounded)","id":"relation-to-published-work-honestly-bounded"},{"depth":2,"text":"v2 — The refined law: the family exponent (2026-07-02, later the same day)","id":"v2-the-refined-law-the-family-exponent-2026-07-02-later-the-same-day"},{"depth":2,"text":"What would kill MRH (registered, in the typed-claim-shape sense)","id":"what-would-kill-mrh-registered-in-the-typed-claim-shape-sense"},{"depth":2,"text":"Practical consequence (the part a lab uses tomorrow)","id":"practical-consequence-the-part-a-lab-uses-tomorrow"}],"html":"<h1 id=\"the-monotone-reparameterization-hypothesis-mrh\">The Monotone Reparameterization Hypothesis (MRH)</h1><blockquote>\n<p>Status: unified statement, 2026-07-02. Every numeric claim below is drawn\nfrom provenance-hashed artifacts in <code>data/y_matrix_runs/</code> and sealed as\ntype-checked Lean modules (<code>lean/Round1..6_*.lean</code>). Novelty claims are\nbounded by two adversarially-verified deep-research sweeps (≈200 sources,\n3-vote verification); community replication is the final arbiter.</p>\n</blockquote>\n<h2 id=\"the-claim-in-one-sentence\">The claim, in one sentence</h2><p><strong>Foundation interatomic potentials learn the physics of a property family up\nto a single monotone warp of the property axis: they get the <em>order</em> of\nnature right and the <em>size</em> wrong, in a way that is one invertible distortion\nper (model, property family) — measurable with a handful of anchors,\ninvertible with proof, and provably absent where training data was too thin\nto fix even the order.</strong></p>\n<p>In plain language: ask a foundation model <em>which</em> metal has the higher\nsurface energy and it almost never lies; ask it <em>how much</em> and it is wrong by\na warp — the same warp for every metal and every facet in the family. Warps\ncan be measured and undone. Scrambles cannot — and we can prove, in a proof\nchecker, which one you have.</p>\n<h2 id=\"formal-statement\">Formal statement</h2><p>For model M and property family F, there exists a strictly increasing map\nw_{M,F} such that predictions satisfy P ≈ w_{M,F}(T) across materials, where\nT is the reference value. Corollaries:</p>\n<ol>\n<li><strong>Ordinal invariance</strong> — rankings of T are preserved by P (the\ninvertibility condition for w).</li>\n<li><strong>Family-level warp</strong> — one w serves all properties in F (facets of a\nsurface share w; the family, not the property, is the warp&#39;s domain).</li>\n<li><strong>Training-distribution control</strong> — w&#39;s distance from identity is set by\ntraining-data bias, not architecture; better-distributed retraining moves\nw toward identity without changing what is ordered.</li>\n<li><strong>Correction = inversion</strong> — calibration is w⁻¹, computable from a few\nanchor (P, T) pairs by isotonic regression; it lives in the monotone\ngroup.</li>\n<li><strong>The boundary</strong> — where training data was insufficient to fix even the\norder, no monotone correction exists (provably), and the model&#39;s output\nfor that (M, F) must be discarded, not calibrated.</li>\n</ol>\n<h2 id=\"evidence-all-measured-on-the-21-material-4-model-y-matrix-this-repo\">Evidence (all measured on the 21-material × 4-model Y-matrix, this repo)</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>#</th>\n<th>Observation</th>\n<th>Where sealed</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"#\">1</td>\n<td data-label=\"Observation\">Rankings survive where magnitudes fail: ρ = 0.82–1.00 across surfaces, vacancies, B₀; 22/22 facet orderings; MPA-0&#39;s γ₁₁₁ ranking is <em>exactly</em> the reference permutation</td>\n<td data-label=\"Where sealed\"><code>Round3_Ordinal.lean</code> (23 thms — kernel checks the order itself)</td>\n</tr>\n<tr>\n<td data-label=\"#\">2</td>\n<td data-label=\"Observation\">Scalar (affine) corrections fail at every grain — per-model, per-material, within-family</td>\n<td data-label=\"Where sealed\"><code>Round1_H4.lean</code>; H4 artifact</td>\n</tr>\n<tr>\n<td data-label=\"#\">3</td>\n<td data-label=\"Observation\">Monotone correction succeeds where MRH says it must: CHGNet surfaces 36→9%, MPA-0 γ₁₁₁ 9→2% (LOO); the correction computed in-kernel for the flagship (Pt: 12.9→0.6%)</td>\n<td data-label=\"Where sealed\"><code>Round4_Isotonic.lean</code></td>\n</tr>\n<tr>\n<td data-label=\"#\">4</td>\n<td data-label=\"Observation\"><strong>Family-level warp (confirmed prediction):</strong> γ₁₀₀-fitted map corrects unseen γ₁₁₀/γ₁₁₁ — CHGNet 26.8→9.9%, MACE-small 13.0→6.7% — <em>beating same-facet calibration</em></td>\n<td data-label=\"Where sealed\"><code>Round6_MRH.lean</code></td>\n</tr>\n<tr>\n<td data-label=\"#\">5</td>\n<td data-label=\"Observation\"><strong>Warp→identity ordering (confirmed prediction):</strong> pooled warp magnitude 0.051 &lt; 0.099 &lt; 0.120 &lt; 0.388 strictly orders MPA-0 &lt; MACE-med &lt; MACE-small &lt; CHGNet by training lineage</td>\n<td data-label=\"Where sealed\"><code>Round6_MRH.lean</code></td>\n</tr>\n<tr>\n<td data-label=\"#\">6</td>\n<td data-label=\"Observation\">Bias/variance split: OMat retraining removes warp bias (s 1.08→0.97, Fe flips soft→stiff) but not per-material variance</td>\n<td data-label=\"Where sealed\">R2 artifact; <code>Round2_Verdicts.lean</code></td>\n</tr>\n<tr>\n<td data-label=\"#\">7</td>\n<td data-label=\"Observation\">The boundary is real and provable: MPtrj SFE rank inversion ⇒ <strong>no monotone correction exists</strong> (quantified impossibility theorem, not policy)</td>\n<td data-label=\"Where sealed\"><code>Round4_Isotonic.lean</code> (<code>sfe_mptrj_uncorrectable</code>)</td>\n</tr>\n<tr>\n<td data-label=\"#\">8</td>\n<td data-label=\"Observation\">The warp is a <em>single-model</em> object: separately-fitted potential families (EAM, 5 and 24-cell tests) do not share a warp and are not family-calibratable</td>\n<td data-label=\"Where sealed\">this session&#39;s EAM artifacts</td>\n</tr>\n</tbody></table></div><h2 id=\"relation-to-the-hyper-ribbon-program\">Relation to the hyper-ribbon program</h2><p>Round-1&#39;s registered H1/H2 &quot;kills&quot; refuted <strong>linear</strong> low-dimensionality\n(participation ratio, leading-mode cosines — linear instruments). MRH locates\nthe structure those instruments could not see: the error manifold is\nlow-dimensional <strong>after allowing monotone reparameterization</strong> — one curved\ndegree of freedom per family. <strong>The ribbon is real, and it is curved.</strong> The\nTranstrum–Sethna sloppiness picture survives in nonlinear form; the linear\nkills were the necessary step that forced the nonlinear formulation.</p>\n<h2 id=\"relation-to-published-work-honestly-bounded\">Relation to published work (honestly bounded)</h2><ul>\n<li>PES softening (Deng et al. 2024/25) documents systematic <em>underprediction</em>\nand corrects with per-system linear rescaling. MRH subsumes it: softening\nis the bias component of a monotone warp; our data show the warp is\nneither linear nor per-system-scalar (H4, evidence #2) but family-monotone\n(#4).</li>\n<li>Our two verified literature sweeps found <strong>no published\ntransferability-of-correction study and no ordinal/cardinal decomposition\nof foundation-model error</strong> — the field&#39;s own open questions list asks for\nexactly these. To the limit of that search, evidence #1–5 and the\nimpossibility-gate methodology (#7) are new.</li>\n<li>Machine-checked certificates for benchmark claims, corrections computed in\nthe proof kernel, and compile-time applicability gates have no precedent\nwe could find in any materials-simulation toolchain.</li>\n</ul>\n<h2 id=\"v2-the-refined-law-the-family-exponent-2026-07-02-later-the-same-day\">v2 — The refined law: the family exponent (2026-07-02, later the same day)</h2><p>Exploration past the nonparametric statement found the warp&#39;s <strong>form</strong>:</p>\n<blockquote>\n<p><strong>pred ≈ c · T^α</strong> — log-affine, R² = 0.93–0.98 (surfaces, B₀).\n<strong>The exponent α belongs to the property family; the prefactor c belongs to\nthe model.</strong></p>\n</blockquote>\n<p>Evidence (sealed in <code>Round7_FamilyExponent.lean</code>; CIs in the analysis run):</p>\n<ul>\n<li>α clusters by family across all four models: surfaces ≈ 1.10 (bootstrap CIs\n1.07–1.22 / 1.06–1.19 / 1.00–1.13 / 1.05–1.15, all overlapping), B₀ ≈ 0.95,\nvacancies ≈ 0.87 (wider CIs). Every model&#39;s surface exponent exceeds every\nmodel&#39;s vacancy and B₀ exponent (point facts, kernel-checked).</li>\n<li>Training moves c toward 1 and barely moves α: CHGNet surfaces c = 0.66 →\nMPA-0 c = 0.98, while α stays ~1.1. This <em>is</em> Round 2&#39;s bias/variance split,\nnow with named parameters.</li>\n<li>The two-parameter correction <strong>beats 8-knot isotonic</strong> LOO on the softened\nmodels (CHGNet 10.0% vs 11.2%; MACE-small 7.45% vs 7.79%) — the minimal\nform is the better estimator.</li>\n<li><strong>One anchor suffices</strong>: family α from <em>other</em> models + a single measured\n(P, T) pair halves CHGNet&#39;s surface error (27.96% → 14.3%).</li>\n</ul>\n<p>Every earlier result is a corollary: scalar corrections fix c while α ≠ 1\n(H4&#39;s failure); monotone maps contain power laws (isotonic&#39;s success); facets\nshare both numbers (family transfer); rank preservation is what any power law\ndoes; and the SFE impossibility marks where even the power-law form collapsed.</p>\n<p>Honest bounds: three families, four models, one lab; vacancy CIs are wide\n(CHGNet&#39;s spans 0.44–1.19); α family-constancy is CI-overlap, not proven\nidentity; MPA-0&#39;s surfaces remain a PASS cell (raw beats every correction).</p>\n<h2 id=\"what-would-kill-mrh-registered-in-the-typed-claim-shape-sense\">What would kill MRH (registered, in the typed-claim-shape sense)</h2><ul>\n<li>A property family with high rank-fidelity whose fitted warp <em>fails</em> to\ntransfer within-family on new materials (hcp metals, alloys are queued).</li>\n<li>Warp non-monotonicity appearing at interpolation densities where ranks are\npreserved (would break corollary 4).</li>\n<li>A model whose warp distance orders <em>against</em> its training-data quality.</li>\n</ul>\n<h2 id=\"practical-consequence-the-part-a-lab-uses-tomorrow\">Practical consequence (the part a lab uses tomorrow)</h2><ul>\n<li>Screening decisions (which candidate is best) are already trustworthy at\nρ ≥ 0.9 — no correction needed; we certify the ordering itself.</li>\n<li>Certification decisions (what is the number) need k anchor measurements\nper (model, family) — k ≈ 8 sufficed here — then w⁻¹ delivers up to 4×\nerror reduction at zero additional simulation, inside existing LAMMPS\nworkflows via log post-processing.</li>\n<li>Where no correction can exist, the system refuses with a proof — before a\ncore-hour is spent. (<code>Round5_LammpsValue.lean</code> holds the cost-accuracy\nfacts: a classical potential beating CHGNet 20/24 cells at ~6× less\ncompute, no GPU.)</li>\n</ul>\n"}