{"id":"conjecture-ledger","title":"The Hypothesis Ledger","subtitle":"Every claim we have tested, its lifecycle status, and why it moved.","category":"conjectures","tags":["ledger","index","overview"],"source":"articles/docs/conjectures/ledger.md","lang":"en","words":532,"readMinutes":2,"toc":[{"depth":2,"text":"The ledger","id":"the-ledger"},{"depth":2,"text":"2026-06-07 Kimi MLIP Import","id":"2026-06-07-kimi-mlip-import"},{"depth":2,"text":"Why this shelf exists","id":"why-this-shelf-exists"}],"html":"<h1 id=\"conjectures-amp-proofs-the-hypothesis-ledger\">Conjectures &amp; Proofs — The Hypothesis Ledger</h1><p>Every claim Lupine has seriously tested, where it stands, and <em>why it moved</em>. This is\nthe structured counterpart to the narrative changelog: the changelog tells the story,\nthe ledger tells the state.</p>\n<p>Each hypothesis is its own entry with a lifecycle <strong>status</strong>. The status legend:</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Status</th>\n<th>Meaning</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Status\"><strong>Supported</strong></td>\n<td data-label=\"Meaning\">Survives the strongest test we have applied so far.</td>\n</tr>\n<tr>\n<td data-label=\"Status\"><strong>Open</strong></td>\n<td data-label=\"Meaning\">Live; evidence is partial or mixed.</td>\n</tr>\n<tr>\n<td data-label=\"Status\"><strong>Refuted by us</strong></td>\n<td data-label=\"Meaning\">We tried to confirm it and the effect did not survive a fair test. The confounder is named.</td>\n</tr>\n<tr>\n<td data-label=\"Status\"><strong>Self-corrected</strong></td>\n<td data-label=\"Meaning\">We <em>announced</em> something, then found our own error and retracted it.</td>\n</tr>\n<tr>\n<td data-label=\"Status\"><strong>Proven (Lean)</strong></td>\n<td data-label=\"Meaning\">Cross-checked by a machine-checked Lean 4 theorem.</td>\n</tr>\n</tbody></table></div><h2 id=\"the-ledger\">The ledger</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Hypothesis</th>\n<th>Status</th>\n<th>One-line resolution</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Hypothesis\">Hyper-ribbon universality (classical potentials)</td>\n<td data-label=\"Status\">Supported · Proven</td>\n<td data-label=\"One-line resolution\">Error vectors occupy a low-dimensional manifold; Lean-grounded.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Hyper-ribbon transfers classical → MLIP</td>\n<td data-label=\"Status\">Under re-audit (2026-06-11)</td>\n<td data-label=\"One-line resolution\">Prior evidence: 14/15 on-ribbon under MACE/CHGNet/Orb-v3. Born screening (replication/error-geometry) excludes 7/45 foundation-model tensors (incl. CHGNet-Fe, MACE-V, Orb Al/Nb/Pb/Pt); per-element counts must be recomputed on screened inputs. Directional structure independently confirmed at n=8–11 models (rank-1 share 0.56–0.94).</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Projected hyper-ribbon release</td>\n<td data-label=\"Status\">Open</td>\n<td data-label=\"One-line resolution\">New Lean-first release lane: prove projected-ribbon gates, then require replay plus cloud evidence before promotion.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Cross-MLIP orthogonal error modes</td>\n<td data-label=\"Status\">Supported</td>\n<td data-label=\"One-line resolution\">MACE and CHGNet have orthogonal error directions on Ag/Nb/Pd.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Au escapes the ribbon under foundation MLIPs</td>\n<td data-label=\"Status\">Open</td>\n<td data-label=\"One-line resolution\">Confirmed for MACE+CHGNet; Ag escape refuted.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Fe magnetic MLIP failure mode</td>\n<td data-label=\"Status\">Open · under re-audit</td>\n<td data-label=\"One-line resolution\">The old PR &gt; 2 trio claim is frozen after Born screening excluded CHGNet-Fe; Fe remains a magnetic/mechanical-stability failure-mode target.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">D-band controls error correlation</td>\n<td data-label=\"Status\"><strong>Refuted by us</strong></td>\n<td data-label=\"One-line resolution\">Sample-size confounder (full-sample ρ = −0.02).</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">MEAM is intrinsically 2-D</td>\n<td data-label=\"Status\"><strong>Refuted by us</strong></td>\n<td data-label=\"One-line resolution\">Matched-n bootstrap: MEAM overlaps Tersoff.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">BCC/FCC &quot;causal shield&quot;</td>\n<td data-label=\"Status\"><strong>Self-corrected</strong></td>\n<td data-label=\"One-line resolution\">The dramatic r 0.90 vs 0.04 was 1.5 % data contamination.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Simpson&#39;s paradox in BCC elastic constants</td>\n<td data-label=\"Status\"><strong>Refuted by us · Lean</strong></td>\n<td data-label=\"One-line resolution\"><code>noSimpsonsInBccEam</code>: the causal graph has no bypass.</td>\n</tr>\n</tbody></table></div><h2 id=\"2026-06-07-kimi-mlip-import\">2026-06-07 Kimi MLIP Import</h2><div class=\"table-wrap\"><table><thead><tr>\n<th>Hypothesis</th>\n<th>Status</th>\n<th>One-line resolution</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Hypothesis\">Parameter-basis Vandermonde decay in foundation MLIPs</td>\n<td data-label=\"Status\"><strong>Refuted by us</strong></td>\n<td data-label=\"One-line resolution\">The 4-model Fisher sweep fails rho &gt;= 1.5; MACE parameter-basis rho is flat while CHGNet/SchNet only reach about rho 0.4.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">MACE irrep-basis Vandermonde threshold</td>\n<td data-label=\"Status\"><strong>Refuted by us</strong></td>\n<td data-label=\"One-line resolution\">Irrep coefficients show real geometric decay (rho 0.3865, R2 0.9807), but still fail the pre-registered rho &gt;= 1.5 threshold.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Weak-form acceleration/refusal theorem</td>\n<td data-label=\"Status\">Open</td>\n<td data-label=\"One-line resolution\">Scalar weak-form gate now builds in Lean; full Lipschitz/reach formalization and deeper-model runtime evidence remain open.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Layerwise distance as MLIP/MD uncertainty signal</td>\n<td data-label=\"Status\">Open</td>\n<td data-label=\"One-line resolution\">Layer-0 distance correlates with force error and helps mixed-reference refusal, but Cu-only reference tuning failed and force-calibrated follow-up is needed.</td>\n</tr>\n<tr>\n<td data-label=\"Hypothesis\">Kimi Cloud Run cross-MLIP v7</td>\n<td data-label=\"Status\">Supported</td>\n<td data-label=\"One-line resolution\">45 MACE/CHGNet/SevenNet elastic calculations are preserved; Fe is a MACE-disagreement sentinel, while Ta/V/Pt have the highest 3-MLIP PR values.</td>\n</tr>\n</tbody></table></div><p>See also:\n<a href=\"../science/kimi-mlip-universality-import.md\" target=\"_blank\" rel=\"noopener\" class=\"ll-raw-source\">Kimi MLIP Universality Import</a> ·\n<a href=\"../runbooks/cross-mlip-cloud-experiment.md\" target=\"_blank\" rel=\"noopener\" class=\"ll-raw-source\">Cross-MLIP Cloud Experiment Runbook</a>.</p>\n<h2 id=\"why-this-shelf-exists\">Why this shelf exists</h2><p>The most defensible thing Lupine produces is not a single result — it is a <em>method that\ncatches its own mistakes</em>. The d-band and MEAM refutations and the BCC/FCC\nself-correction all came from the same matched-n / contamination-gate discipline. Making\nthe refutations as visible as the confirmations is the point: a corpus you can trust is\none that publishes what it killed.</p>\n<p>See also: <a href=\"#/read/formal-proof-ledger\">Formal Proof Ledger</a> ·\n<a href=\"#/read/methodology\">Methodology</a> · <a href=\"#/read/data-provenance\">Data &amp; Provenance</a>.</p>\n"}