{"id":"changelog","title":"Changelog & Progress","subtitle":"What changed, why, results, and suggested next steps. Read this first.","category":"changelog","tags":["changelog","progress","narrative"],"source":"articles/CHANGELOG.md","lang":"en","words":10325,"readMinutes":47,"toc":[{"depth":2,"text":"2026-07-17 - Round-3 consequence wave: B0 correction gate denial + assumption-link CI + contract compiler","id":"2026-07-17-round-3-consequence-wave-b0-correction-gate-denial-assumption-link-ci-contract-compiler"},{"depth":2,"text":"2026-07-16 - Empirical-formal bridge integration wave: contracts, evidence bundles, anti-laundering CI, and UniversalCorrection registry","id":"2026-07-16-empirical-formal-bridge-integration-wave-contracts-evidence-bundles-anti-laundering-ci-and-universalcorrection-registry"},{"depth":2,"text":"2026-07-13 - CI unblock: disk-quota fix for evidence-index workflows + diff-hygiene guard for LAMMPS work dirs","id":"2026-07-13-ci-unblock-disk-quota-fix-for-evidence-index-workflows-diff-hygiene-guard-for-lammps-work-dirs"},{"depth":2,"text":"2026-07-13 - Non-CO₂ environmental-concern contract module","id":"2026-07-13-non-co-environmental-concern-contract-module"},{"depth":2,"text":"2026-07-11 - The anchor-identification laws: what three anchors determine — certified brackets, margin-certified ranking, and refusal completeness","id":"2026-07-11-the-anchor-identification-laws-what-three-anchors-determine-certified-brackets-margin-certified-ranking-and-refusal-completeness"},{"depth":2,"text":"2026-07-11 - Climate-portfolio contract module in Lean","id":"2026-07-11-climate-portfolio-contract-module-in-lean"},{"depth":2,"text":"2026-07-11 - Production wiring of the orphaned certificate gates + rocksalt/halide layout + climate-series Python mirror","id":"2026-07-11-production-wiring-of-the-orphaned-certificate-gates-rocksalt-halide-layout-climate-series-python-mirror"},{"depth":2,"text":"2026-07-10 - Diamond anchor + run-time certificate gate: Si joins the corpus, refusals now block correction inside the policy engine","id":"2026-07-10-diamond-anchor-run-time-certificate-gate-si-joins-the-corpus-refusals-now-block-correction-inside-the-policy-engine"},{"depth":2,"text":"2026-07-10 - Bcc anchor set + flywheel telemetry: refractory metals join the measured-fields corpus, certificates ride the live promotion spans","id":"2026-07-10-bcc-anchor-set-flywheel-telemetry-refractory-metals-join-the-measured-fields-corpus-certificates-ride-the-live-promotion-spans"},{"depth":2,"text":"2026-07-10 - Measured fields rung: Y-matrix corpus bound into per-cell Lean field instances + certificate-carrying promotion gate","id":"2026-07-10-measured-fields-rung-y-matrix-corpus-bound-into-per-cell-lean-field-instances-certificate-carrying-promotion-gate"},{"depth":2,"text":"2026-07-10 - Climate-series physics layer: 73 new Lean theorems for the five material classes","id":"2026-07-10-climate-series-physics-layer-73-new-lean-theorems-for-the-five-material-classes"},{"depth":2,"text":"2026-07-07 - Evidence index: full corpus coverage, measured retrieval, one embedding space","id":"2026-07-07-evidence-index-full-corpus-coverage-measured-retrieval-one-embedding-space"},{"depth":2,"text":"2026-06-19 - LUPI 0.3 Studio, molecule trust, and public-surface split prep","id":"2026-06-19-lupi-0-3-studio-molecule-trust-and-public-surface-split-prep"},{"depth":2,"text":"2026-06-16 — Academic review of the Projection Law / IMMI suite, first fix pass","id":"2026-06-16-academic-review-of-the-projection-law-immi-suite-first-fix-pass"},{"depth":2,"text":"2026-06-14 - LUPI controls palette rollout","id":"2026-06-14-lupi-controls-palette-rollout"},{"depth":2,"text":"2026-06-12 — Repo consolidation and onboarding sprint","id":"2026-06-12-repo-consolidation-and-onboarding-sprint"},{"depth":2,"text":"2026-06-02 — CORRECTION: retracting / bounding the day's ribbon overclaims","id":"2026-06-02-correction-retracting-bounding-the-day-s-ribbon-overclaims"},{"depth":2,"text":"2026-06-02 — Back to the keystone paper: the category error, and the first test of A6","id":"2026-06-02-back-to-the-keystone-paper-the-category-error-and-the-first-test-of-a6"},{"depth":2,"text":"2026-06-02 — Live 3-tier sim campaign: distill is energy-only + per-backend-policy-gated","id":"2026-06-02-live-3-tier-sim-campaign-distill-is-energy-only-per-backend-policy-gated"},{"depth":2,"text":"2026-06-02 — Theorist deep-tier model upgrade: MiniMax M2.7 → M3, gated on a measured A/B","id":"2026-06-02-theorist-deep-tier-model-upgrade-minimax-m2-7-m3-gated-on-a-measured-a-b"},{"depth":2,"text":"2026-05-29 — Neural-symbolic loop: GPU MLIP curvature → machine-checked Lean (0 sorry)","id":"2026-05-29-neural-symbolic-loop-gpu-mlip-curvature-machine-checked-lean-0-sorry"},{"depth":2,"text":"2026-05-29 — Local GPU proof: TorchSim → distill → uplift → formal gate (Ni FCC)","id":"2026-05-29-local-gpu-proof-torchsim-distill-uplift-formal-gate-ni-fcc"},{"depth":2,"text":"2026-05-29 — ATLAS-Lean integration: formal foundations + closed-loop scaffolding","id":"2026-05-29-atlas-lean-integration-formal-foundations-closed-loop-scaffolding"},{"depth":2,"text":"2026-05-18 — Fix mislabeled home-page working-paper banner","id":"2026-05-18-fix-mislabeled-home-page-working-paper-banner"},{"depth":2,"text":"2026-05-19 — paper-build auto-dispatches the Library deploy","id":"2026-05-19-paper-build-auto-dispatches-the-library-deploy"},{"depth":2,"text":"2026-05-18 — Opt-in CI to rebuild the working-paper PDF","id":"2026-05-18-opt-in-ci-to-rebuild-the-working-paper-pdf"},{"depth":2,"text":"2026-05-18 — Fix the broken working-paper PDF","id":"2026-05-18-fix-the-broken-working-paper-pdf"},{"depth":2,"text":"2026-05-18 — Remove Entity Graph; fix callout/filter alignment","id":"2026-05-18-remove-entity-graph-fix-callout-filter-alignment"},{"depth":2,"text":"2026-05-18 — References & Lineage shelf","id":"2026-05-18-references-lineage-shelf"},{"depth":2,"text":"2026-05-18 — Phase 2b: reader-side status filter","id":"2026-05-18-phase-2b-reader-side-status-filter"},{"depth":2,"text":"2026-05-18 — Phase 2: the corpus becomes a ledger (Tier 1 + Tier 2)","id":"2026-05-18-phase-2-the-corpus-becomes-a-ledger-tier-1-tier-2"},{"depth":2,"text":"2026-05-18 — Phase 1b: the deploy was green but the site never changed","id":"2026-05-18-phase-1b-the-deploy-was-green-but-the-site-never-changed"},{"depth":2,"text":"2026-05-18 — Phase 1: unbreak the Library deploy path (self-healing SW)","id":"2026-05-18-phase-1-unbreak-the-library-deploy-path-self-healing-sw"},{"depth":2,"text":"2026-05-18 — Revive the Library as the public thinking surface","id":"2026-05-18-revive-the-library-as-the-public-thinking-surface"},{"depth":2,"text":"2026-05-16 — Wire Phoenix evals end to end","id":"2026-05-16-wire-phoenix-evals-end-to-end"},{"depth":2,"text":"2026-05-16 — De-myopize the corpus (the hyper-ribbon is not an artifact)","id":"2026-05-16-de-myopize-the-corpus-the-hyper-ribbon-is-not-an-artifact"},{"depth":2,"text":"2026-05-16 — Self-correction: the BCC/FCC \"causal shield\" was contamination","id":"2026-05-16-self-correction-the-bcc-fcc-causal-shield-was-contamination"},{"depth":2,"text":"2026-05-17 — Phase D: close the loop with real physics","id":"2026-05-17-phase-d-close-the-loop-with-real-physics"},{"depth":2,"text":"2026-05-17 — The self-improving eval loop (Evolver spine)","id":"2026-05-17-the-self-improving-eval-loop-evolver-spine"},{"depth":2,"text":"2026-05-04…05 — Hypothesis closures: the corpus refutes itself, correctly","id":"2026-05-04-05-hypothesis-closures-the-corpus-refutes-itself-correctly"},{"depth":2,"text":"2026-05-15…16 — One LLM path, eval-aware routing, AI Gateway","id":"2026-05-15-16-one-llm-path-eval-aware-routing-ai-gateway"},{"depth":2,"text":"2026-05-06…18 — atlas-view: streaming, render polish, curated gallery","id":"2026-05-06-18-atlas-view-streaming-render-polish-curated-gallery"},{"depth":2,"text":"How to add an entry","id":"how-to-add-an-entry"}],"html":"<h1 id=\"changelog-amp-progress\">Changelog &amp; Progress</h1><p>A working log of what changed in the Lupine research program — not just <em>what</em> we did, but\n<strong>why</strong>, what the <strong>results</strong> were, and the <strong>suggested next steps</strong> to extract more value.\nThis is meant to be read: it is the narrative spine of the corpus, and it is published on\n<a href=\"https://library.lupine.science\">library.lupine.science</a>.</p>\n<p>Format for every entry:</p>\n<ul>\n<li><strong>Why</strong> — the problem or question that motivated the change.</li>\n<li><strong>What</strong> — what we actually did.</li>\n<li><strong>Results</strong> — what we observed, including null and negative results.</li>\n<li><strong>Next</strong> — the highest-value follow-up, so the log compounds instead of just accumulating.</li>\n</ul>\n<p>Newest first. Dates are absolute.</p>\n<hr>\n<h2 id=\"2026-07-17-round-3-consequence-wave-b0-correction-gate-denial-assumption-link-ci-contract-compiler\">2026-07-17 - Round-3 consequence wave: B0 correction gate denial + assumption-link CI + contract compiler</h2><ul>\n<li><strong>Why.</strong> Round-3 produced a kill-condition verdict for B0 correction (both\nrocksalt and perovskite worsened under the frozen rule) and a confirmed\nscope for a0. The formal plane needed to encode that disposition as a\nfail-closed gate, and the CI plane needed to verify every theorem and\npremise has a receipt.</li>\n<li><strong>What.</strong><ul>\n<li>Lean <code>UniversalCorrection/Empirical/Registry.lean</code> gains\n<code>CorrectionGatePolicy</code> with <code>B0 -&gt; deny/contradicting_evidence</code> and\n<code>a0 -&gt; allow/scope_matched_same_class_a0</code>.</li>\n<li><code>registry/claims/correction.b0.v1.json</code> and\n<code>registry/claims/fcc.b0.anticorrelation.v1.json</code> rewritten to describe\ncontradicting evidence rather than licensed improvement.</li>\n<li><code>data/discovery_gates/licenses.v1.json</code> gains a <code>correction_gate</code> deny\nfor B0.</li>\n<li><code>tools/check_assumption_links.py</code> + <code>assumption-links</code> verify job:\nassurance-exported theorems must have <code>@[requires(contract_id)]</code>,\npremises must have valid bundle references or explicit unsupported tags.</li>\n<li><code>tools/atlas_theorem_sync.py --compile-gates</code> compiles locked\nClaimContracts into a provenance-bearing runtime gate manifest.</li>\n<li><code>python/tests/test_d1_migrations.py</code> updated to expect the intentional\n<code>0011_atlas_schema_reconciliation.sql</code> reapplication guard.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (3732 jobs). 115 focused B0/license/Round-3\nPython tests pass; 14 assumption-link tests pass; 56 evidence-formal bridge\ntests pass.</li>\n<li><strong>Next.</strong> Finish <code>t_64d423d6</code> atlas-sync contract compiler (EvidenceBundle\nresolution, duplicate gate-ID guard, premise-policy semantics), then run\nthe end-to-end QA replay (<code>t_163d358b</code>).</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-16-empirical-formal-bridge-integration-wave-contracts-evidence-bundles-anti-laundering-ci-and-universalcorrection-registry\">2026-07-16 - Empirical-formal bridge integration wave: contracts, evidence bundles, anti-laundering CI, and UniversalCorrection registry</h2><ul>\n<li><strong>Why.</strong> The Round-3 empirical results (B0 refuted, a0 confirmed) were not yet\nlocked into the formal evidence plane. There was no machine-readable registry\nlinking Lean theorems to their empirical scope, no CI guard preventing\nassumption laundering (status upgrades without new evidence), and no D1\nschema for versioned claim/evidence contracts.</li>\n<li><strong>What.</strong><ul>\n<li>Added versioned schemas for EvidenceBundle, ClaimContract, and\nCampaignManifest; backfilled the three Round-3 claim contracts and their\ncontent-addressed evidence bundles.</li>\n<li>Added <code>tools/generate_assumptions.py</code> to derive\n<code>registry/assumptions.v1.json</code> and <code>registry/snapshots/current.lock.json</code>\nfrom the contracts/bundles.</li>\n<li>Added D1 migration <code>0011_claim_evidence_contracts.sql</code> with append-only\n<code>status_event</code> and <code>evidence_bundle_id</code> FK so D1 status changes can be\ntied to receipts.</li>\n<li>Hardened legacy migration <code>0004_claims.sql</code> so it applies cleanly on both\nfresh and production-shaped databases.</li>\n<li>Added anti-laundering CI (<code>tools/check_status_evidence.py</code> +\n<code>.github/workflows/anti-laundering.yml</code>) that rejects status/assurance\nchanges unless a new EvidenceBundle hash appears.</li>\n<li>Added <code>OpenDistillationFactory.UniversalCorrection.Empirical.Registry</code>\n(172 theorem inventory entries across 17 active correction modules) plus\n<code>ActiveSampling.lean</code>, <code>SubspaceCorrectionScheme.lean</code>, and\n<code>UniversalCorrectionScheme.lean</code>. Excluded quarantined\n<code>AlloyResidualTransfer</code> from the registry.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (3675 jobs, 0 <code>sorry</code>). Python contract tests\n(40) pass, <code>tools/generate_assumptions.py --check</code> passes, and the\nanti-laundering checker accepts/rejects as expected.</li>\n<li><strong>Next.</strong> Finish the three running kanban tasks (<code>t_9d49164c</code>,\n<code>t_223ee397</code>, <code>t_64d423d6</code>) and then run the end-to-end QA integration test\n(<code>t_163d358b</code>) to prove Round-3 propagates with zero manual touches.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-13-ci-unblock-disk-quota-fix-for-evidence-index-workflows-diff-hygiene-guard-for-lammps-work-dirs\">2026-07-13 - CI unblock: disk-quota fix for evidence-index workflows + diff-hygiene guard for LAMMPS work dirs</h2><ul>\n<li><strong>Why.</strong> Two recurring CI failures were blocking the research pipeline:\nthe <code>Evidence Index</code> and <code>Evidence Index Nightly</code> workflows were exceeding\nthe GitHub runner disk quota because <code>sentence-transformers</code> pulled the\nfull CUDA <code>torch</code> distribution; and the <code>Verify</code> diff-whitespace check was\nfailing on generated LAMMPS log files inside <code>data/lammps_validation/**/work/</code>.</li>\n<li><strong>What.</strong><ul>\n<li>Added an explicit CPU-only <code>torch</code> install step to\n<code>.github/workflows/evidence-index.yml</code> and\n<code>.github/workflows/evidence-nightly.yml</code> before <code>pip install -e .</code>, so\n<code>sentence-transformers</code> reuses the CPU wheel instead of downloading\nseveral gigabytes of NVIDIA CUDA packages.</li>\n<li>Added <code>data/lammps_validation/**/work/</code> to <code>.gitignore</code> so generated\nLAMMPS run artifacts (which contain trailing whitespace by design) are\nexcluded from the diff-hygiene gate.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> Local reproduction of the evidence-index unit tests now passes\nafter installing CPU torch; the diff-hygiene rule now ignores LAMMPS work\ndirectories. The PR branch <code>wip/statics-discovery-gates</code> still needs its\ncommitted <code>work/</code> directory removed and rebased to pick up the <code>.gitignore</code>\nchange.</li>\n<li><strong>Next.</strong> Clean the committed <code>data/lammps_validation/2026-07-13-ni-eam-elastic/work/</code>\ndirectory from <code>wip/statics-discovery-gates</code> and re-run <code>Verify</code>.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-13-non-co-environmental-concern-contract-module\">2026-07-13 - Non-CO₂ environmental-concern contract module</h2><ul>\n<li><strong>Why.</strong> The climate-series proof pack and the Lupine Science articles now\ncover methane, HFC refrigerants, water remediation, air remediation,\ncritical-mineral/PFAS circularity, and embodied carbon in cement. The formal\nevidence plane had no build-locked contract that maps these concerns to the\nexisting theory modules and material classes.</li>\n<li><strong>What.</strong><ul>\n<li>Added <code>Validation/NonCO2ClimateConcerns.lean</code> (9 theorems, T259–T267):\nan inductive <code>EnvironmentalConcern</code> for the six non-CO₂ concerns;\nper-concern strings for mechanism, governing theory, correction path, and\ndata status; linkage back to <code>ClimatePortfolio.MaterialClass</code> for the\nfive concerns that map to the existing portfolio; and invariant witnesses\nthat restate existing theorems from <code>SorptionStability</code>,\n<code>BarrierArrhenius</code>, <code>RankingIntegrity</code>, <code>ScalingVolcano</code>, and\n<code>DefectStability</code>.</li>\n<li>Wired the module into <code>Vision.lean</code> with <code>#check</code> locks and bumped the\ntheorem-inventory count from 262 to 271.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (3665 jobs, 0 <code>sorry</code>). Python unit tests\n(67) and Rust check pass. The status board now prints 271 formally proven\ntheorems.</li>\n<li><strong>Next.</strong> Add quantitative abatement envelopes for the non-CO₂ concerns once\nthe underlying statics and lifecycle data are available; bind the first\nrocksalt/halide and sorbent cells to turn the <code>layoutReadyPending</code> concerns\ninto bound portfolio witnesses.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-11-the-anchor-identification-laws-what-three-anchors-determine-certified-brackets-margin-certified-ranking-and-refusal-completeness\">2026-07-11 - The anchor-identification laws: what three anchors determine — certified brackets, margin-certified ranking, and refusal completeness</h2><ul>\n<li><strong>Why.</strong> The measured-fields rung binds anchors into per-cell field\ninstances, and <code>AnchoredField.lean</code> <em>claims in prose</em> that the clamped\nstep field is &quot;the conservative envelope of the measured data … with no\nextrapolation freedom&quot; — but no theorem said what the anchors actually\ndetermine about the unknown true field, when a corrected value or\n<em>ranking</em> may be trusted (the promotion gate&#39;s core operation), or\nwhether the 49 kernel refusals rule out only our step constructor or\nevery softening field.</li>\n<li><strong>What.</strong><ul>\n<li><code>Theory/AnchorBracket.lean</code> (54 theorems, T197–T250): existence ↔\nadmissibility iffs per layout with scaled-integer bridges (<strong>refusal\ncompleteness</strong> — a refused cell is consistent with <em>no</em> <code>ErrorField</code>\nat all); the <strong>one-scalar reduction</strong> — on in-range configurations any\nconsistent field&#39;s sum is the step-field sum plus\n<code>count_gap · (P(gap) − p_lo)</code> with <code>P(gap)</code> forced into the neighboring\nanchor interval (fcc gap c = 10, bcc c = 5, diamond and rocksalt none); deep/shallow\n<strong>envelope extremality</strong>; <strong>certified correction brackets</strong>\n(<code>E_ref ≤ corrected ≤ E_ref + count·width</code>; diamond: exact recovery even\nat the measured tier); the <strong>two-point certification test</strong> — corrected\norder holds for <em>every</em> consistent field iff it holds at the two\nenvelope fields — plus the interval-<strong>separation margin rule</strong>;\n<strong>certified Arrhenius rate caps</strong> (<code>exp(width/kT)</code> factor bounds via the\nexact amplification identity); <strong>honest non-identifiability</strong> (below the\nlowest anchor every value <code>v ≤ p_lo</code> is realized, and at the measured\ntier even the gap is free — the tier-2 axioms are what buy finite\nbrackets); and <code>FieldDomain → InRange</code> glue making explicit that the\ndefault <code>[4, 12]</code> runtime domain does <em>not</em> discharge the fcc bracket\nprecondition (needs <code>cmin ≥ 8</code>).</li>\n<li><code>Validation/AnchorBracketCertificates.lean</code> (8 theorems, T251–T258):\ncorpus-bound certificates — impossibility for the flagship refusals\n(composed <em>through</em> the generated <code>field_refused_*</code> theorems, so corpus\nregeneration desync breaks the build), gap/width certificates\n(chgnet/Ni 537×10⁻⁴ eV/atom, chgnet/Fe 256×10⁻⁴), the cross-model\ncomparison (on Ni, mace-mp-medium&#39;s bracket is provably more than 6×\ntighter than chgnet&#39;s), and chgnet/Si exactness over any consistent\nmeasured field.</li>\n<li>Runtime mirror (<code>lupine_distill.odf.field_certificates</code>): every\n<code>AnchorCertificate</code> now carries the identification payload\n(<code>gap_coordination</code>, <code>bracket_width_scaled</code>, <code>identification_ref</code>), and\n<code>check_bracket_separation</code> mirrors the separation rule (strict\ninequality, budget on the higher candidate only, float-boundary and\nuniform-bias caveats documented). 12 new mirror tests pin the semantics\non the same corpus witnesses the Lean locks use.</li>\n<li><code>Vision.lean</code>: T197–T258 checks, four corpus width <code>#guard</code> locks,\ntheorem count → 262 (merged atop the climate-portfolio push).</li>\n</ul>\n</li>\n<li><strong>Results.</strong> Full <code>lake build</code> green at 3663 jobs, zero <code>sorry</code>, zero new\naxioms (flagships checked: <code>propext</code>/<code>Classical.choice</code>/<code>Quot.sound</code>\nonly); 35/35 field-certificate mirror tests pass; evidence manifest check\ngreen. An independent adversarial audit (vacuity, inequality directions,\nboundary strictness, cast traps, prose-vs-formal, corpus numerals)\nreturned &quot;survives with fixes&quot;; all findings fixed — notably the flagship\ndiamond certificate was strengthened from a degenerate self-instantiation\nto quantification over any consistent measured field.</li>\n<li><strong>Next.</strong> Interval anchors (measurement-noise-robust admissibility and\nbrackets with a certified stability radius); a bcc/diamond mirror of the\nkinetics caps; consuming <code>bracket_width_scaled</code> inside the promotion\ngate&#39;s ranking path so cross-model cell preference (e.g. Ni) is applied,\nnot just certified.</li>\n</ul>\n<h2 id=\"2026-07-11-climate-portfolio-contract-module-in-lean\">2026-07-11 - Climate-portfolio contract module in Lean</h2><ul>\n<li><strong>Why.</strong> The climate-series proof pack makes quantitative claims about five\nmaterial classes, but the formal evidence plane had no single build-locked\ndocument that maps each class to its governing theory, failure mode,\ncorrection path, and current data status. We needed a contract module that\nboth validates the portfolio envelope and records the rocksalt/halide\nlayout-ready / data-pending state.</li>\n<li><strong>What.</strong><ul>\n<li>Added <code>Validation/ClimatePortfolio.lean</code>: an inductive <code>MaterialClass</code> for\nthe five targets (cobalt-free LMR cathodes, halide solid electrolytes,\nMOF DAC sorbents, electrochemical ammonia catalysts, lead-free perovskites);\nper-class strings for governing theory, uMLIP failure mode, and correction\npath; decidable portfolio-envelope certificates; rocksalt-layout-existence\nand halide-unbound-pending-defect-runs certificates; and per-class\nscreening invariants restated as witnesses backed by existing theorems in\n<code>RankingIntegrity</code>, <code>BarrierArrhenius</code>, <code>SorptionStability</code>,\n<code>ScalingVolcano</code>, and <code>DefectStability</code>.</li>\n<li>Wired <code>ClimatePortfolio</code> into <code>Vision.lean</code> with <code>#check</code> locks and bumped\nthe theorem-inventory count from 190 to 200.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (3,662 jobs, 0 <code>sorry</code>). Python unit tests\n(55) all pass. The status board now prints 200 formally proven theorems.</li>\n<li><strong>Next.</strong> Bind the first rocksalt/halide cells once charge-balanced slab +\n  vacancy formation runs are available for MgO/NaCl/Li₂ZrCl₆/Li₃YCl₆; then\n  extend <code>ClimatePortfolio</code> with measured-field witnesses for the halide\n  electrolyte class.</li>\n</ul>\n<h2 id=\"2026-07-11-production-wiring-of-the-orphaned-certificate-gates-rocksalt-halide-layout-climate-series-python-mirror\">2026-07-11 - Production wiring of the orphaned certificate gates + rocksalt/halide layout + climate-series Python mirror</h2><ul>\n<li><strong>Why.</strong> The climate-series formalization introduced three new certificate\npredicates — field-domain admission, ranking-inversion detection, and\nbarrier-underestimation conservatism — plus a <code>ClimateSeries</code> validation pack\nand the rocksalt/halide family. Only the anchor-admissibility gate was\nactually reaching production: <code>check_field_domain</code>, <code>check_ranking_pair</code>, and\n<code>BarrierArrhenius.softened_barrier_underestimates</code> were tested but never\ncalled by the runtime policy engine or promotion gate. Meanwhile the\nrocksalt cells were recorded as <code>unbound_structures</code> rather than having an\nexplicit layout ready for charge-balanced slab/defect data.</li>\n<li><strong>What.</strong><ul>\n<li><code>Theory/AnchoredField.lean</code>: added the rocksalt/halide layout —\n<code>stepFieldRocksalt</code> (single c=5 anchor, bulk pin at c=6),\n<code>mkAnchoredFieldRocksalt : ErrorField 6</code>, <code>mkMeasuredFieldRocksalt</code>, and\ndecidable <code>scaledAnchorRocksaltValid</code> — plus the evaluation theorems\n(<code>_at_vacancy</code>, <code>_at_bulk</code>, <code>_clamped_below</code>, <code>_toMeasuredField</code>,\n<code>scaledAnchorRocksaltValid_example</code>). The layout is ready for MgO/NaCl and\nthe Li–M–Cl halide electrolytes once their slab/defect observables are\nmeasured. Zero <code>sorry</code>, zero new axioms.</li>\n<li><code>lupine_distill.odf.field_certificates</code>: added <code>BarrierCertificate</code> and\n<code>check_barrier_conservatism</code>, mirroring\n<code>BarrierArrhenius.softened_barrier_underestimates</code> and\n<code>softening_never_hides_conductor</code>; added rocksalt to\n<code>ANCHOR_COORDINATIONS</code> / <code>_ANCHOR_REF_KEYS</code> and the corresponding theorem\nrefs.</li>\n<li><code>lupine_distill.odf.climate_series</code> (new): Python mirror of\n<code>Validation.ClimateSeries</code> with typed certificates for all 10 headline\nclaims (synthesis funnel, A-Lab novelty, kernel-rejected zero margin,\ncorrected strict improvement, blind residuals, Ni/Cu error reductions,\nportfolio envelope, inventory floor), theorem refs, and pass/fail checks.\nInventory-floor defaults and the Lean <code>proof_pack_inventory_floor</code> theorem\nupdated to the current build state: 51 modules, 190 build-locked theorems,\n~640 declarations, zero <code>sorry</code>.</li>\n<li><code>lupine_distill_runtime.policy_engine</code>: <code>_domain_action</code> now calls\n<code>check_field_domain</code> when a prediction or context carries\n<code>first_shell_coordinations</code>; out-of-domain atoms trigger a\n<code>skip_correction</code> action backed by <code>FieldDomain.refusal_has_witness</code> and\nstrip the support model before any correction is applied.</li>\n<li><code>lupine_distill.odf.promotion_gate</code>: added <code>reference_ranking</code> and\n<code>model_ranking</code> metadata fields; <code>evaluate</code> now checks every adjacent pair\nwith <code>check_ranking_pair</code>. A machine-checked inversion downgrades\n<code>promote</code> → <code>review</code> or <code>review</code> → <code>reject</code>, because no monotone\nrecalibration can rescue it.</li>\n<li><code>python/scripts/bind_env_field_instances.py</code>: generalized <code>_bind_cell</code> to\nhandle layouts with no facets and optional vacancy blocks; cells missing\nall bindable observables are skipped with a warning instead of crashing.\nAdded the rocksalt layout (MgO, NaCl, LiCl, Li3YCl6, Li2ZrCl6) and moved\nlayered oxides to <code>unbound_structures</code> with a clear data requirement.\nRegenerated <code>EnvFieldInstances.lean</code> and <code>env_field_binding_report.json</code>;\ncorpus remains 68 cells → 19 instances + 49 refusals, now with a\n<code>rocksalt</code> structure entry at 0 cells.</li>\n<li>Tests: added <code>python/tests/test_climate_series.py</code>,\n<code>python/tests/test_bind_env_field_instances.py</code>, rocksalt/barrier/ranking\ntests in <code>test_field_certificates.py</code>, and domain-gate tests in\n<code>test_certificate_gate.py</code>.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (3,661 jobs, 0 <code>sorry</code>). Python test suite\nnow contains 55 tests, all passing. The field-domain gate, ranking\ngate, and barrier-conservatism certificate are now wired into production\npaths. Rocksalt layout exists but cannot bind until charge-balanced slab +\nvacancy runs are added; the binder documents this explicitly rather than\nfailing silently.</li>\n<li><strong>Next.</strong> Add charge-balanced rocksalt slab and vacancy formation targets +\n  statics runs so the halide electrolyte portfolio target can be formally\n  bound; surface <code>ranking_inverted</code> and <code>field_domain</code> skip events in the\n  Phoenix flywheel dashboards; expose <code>climate_series</code> certificates in the\n  promotion packet renderer and website article footnotes.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-10-diamond-anchor-run-time-certificate-gate-si-joins-the-corpus-refusals-now-block-correction-inside-the-policy-engine\">2026-07-10 - Diamond anchor + run-time certificate gate: Si joins the corpus, refusals now block correction inside the policy engine</h2><ul>\n<li><strong>Why.</strong> Two follow-ups from the bcc rung (below): the remaining Y-matrix\nstructures had no anchor layouts, and refusal certificates reached\ntelemetry but not the run-time correction path — a tier-2-refused cell\nwould still have been corrected if a support model covered it, with the\nrefusal only flagged after the fact.</li>\n<li><strong>What.</strong><ul>\n<li><code>Theory/AnchoredField.lean</code>: the diamond layout — <code>stepFieldDiamond</code>\n(single vacancy anchor: each diamond vacancy exposes 4 first-shell atoms\nat c=3, bulk pin at the diamond coordination c=4),\n<code>mkAnchoredFieldDiamond : ErrorField 4</code> / <code>mkMeasuredFieldDiamond</code>, and\nthe decidable <code>scaledAnchorDiamondValid</code> — 9 new theorems (T182–T190),\nzero sorry, zero new axioms. One anchor is enough for both tiers: the\nmeasured tier unconditionally, the directional tier iff <code>p3 ≤ 0</code>.</li>\n<li><code>bind_env_field_instances.py</code> generalized to variable-anchor layouts\n(targets now also read <code>beyond_metals.json</code>) and bound the 4 Si diamond\ncells. Corpus: 68 cells → 19 instances + 49 refusals.</li>\n<li>The run-time gate: <code>CertificateGate</code> +\n<code>CertificateGatedPolicyEngine</code> in <code>lupine_distill_runtime.policy_engine</code>\nindex the binding report&#39;s refusals (via the shared\n<code>certificates_from_binding_report</code> mirror in\n<code>lupine_distill.odf.field_certificates</code>) and decide refused cells\nWITHOUT the support model — the correction is never applied, not\napplied-then-flagged. Decisions carry a <code>skip_correction</code> action naming\nthe refusal theorem, exact scaled anchors, and corpus hash;\n<code>DistillSession</code> wires the gate by default (&quot;auto&quot;: repo report when\npresent), surfaces it in the runtime summary, and the cell runner\nexposes <code>--env-field-report</code> / <code>MLIP_DISTILL_ENV_FIELD_REPORT</code>. The\nfinite/explosion guards still run on gated cells, and the prediction\nitself is not refused — only its correction is withheld.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> All four Si diamond cells soften monotonically and bind as\ntier-2 instances — uMLIPs underestimate the Si vacancy formation energy by\n1.1–2.8 eV against the DFT-PBE 3.63 eV reference (chgnet worst at 0.87 eV;\nper-atom anchors −2778e-4 to −6906e-4 eV). 190 build-locked theorems, 418\ntheorem declarations, 0 <code>sorry</code>, <code>lake build</code> green. Gate behavior pinned\nby tests: chgnet/Pt (fcc) and chgnet/Ta (bcc) predictions pass through\nuncorrected with the Lean refusal in the decision record; chgnet/Ni still\ncorrects; unbound models are untouched. Notable negative result: the\nrocksalt cells (MgO, NaCl) measure <em>no</em> anchor observables — their statics\nruns carry only EOS + lattice results — so no rocksalt layout is possible\nfrom the existing corpus; the binder now records this as\n<code>unbound_structures</code> in the report instead of silently skipping them.</li>\n<li><strong>Next.</strong> Run charge-balanced rocksalt slab + vacancy statics for MgO/NaCl\nto give the ionic family its anchors, and surface the runtime\n<code>skip_correction</code> events in the Phoenix flywheel dashboards alongside the\npromotion-packet certificate spans.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-10-bcc-anchor-set-flywheel-telemetry-refractory-metals-join-the-measured-fields-corpus-certificates-ride-the-live-promotion-spans\">2026-07-10 - Bcc anchor set + flywheel telemetry: refractory metals join the measured-fields corpus, certificates ride the live promotion spans</h2><ul>\n<li><strong>Why.</strong> The measured-fields rung (below) covered only the 36 fcc cells; the\nbcc refractory metals (Cr/Fe/Mo/Nb/Ta/V/W — 28 more Y-matrix cells with\nDFT-PBE targets already in the corpus) had no anchor layout, and the field\ncertificates stopped at the promotion gate&#39;s metadata instead of reaching\nthe live flywheel telemetry (<code>tools/mlip_local_promotion.py</code> → Phoenix).</li>\n<li><strong>What.</strong><ul>\n<li><code>Theory/AnchoredField.lean</code>: the bcc layout — <code>stepFieldBcc</code> (clamped\nstep interpolation of γ₁₀₀ → c=4, γ₁₁₀ → c=6, E_vac → c=7, bulk pin at\nthe bcc first-shell coordination c=8, the unanchored c=5 gap held at the\ndeeper (100) value, exactly as the fcc field holds c=10),\n<code>mkAnchoredFieldBcc : ErrorField 8</code> / <code>mkMeasuredFieldBcc : MeasuredField 8</code>, and the decidable <code>scaledAnchorsBccValid</code> predicate —\n11 new theorems (T171–T181), zero sorry, zero new axioms.</li>\n<li><code>python/scripts/bind_env_field_instances.py</code>: generalized over structure\nlayouts (per-atom facet areas: fcc a₀²/2 and √3a₀²/4; bcc a₀² and\na₀²/√2; vacancy shells 12 vs 8) and extended to the 28 bcc cells; report\nschema bumped to <code>lupine.env_field_binding_report.v2</code> (per-structure\nlayout + counts, nested anchors). Fcc literals verified byte-identical\nacross the refactor.</li>\n<li><code>lupine_distill.odf.field_certificates</code>: <code>check_anchor_admissibility</code>\ntakes a <code>structure</code> argument, mirrors <code>scaledAnchorsBccValid</code> for bcc\ncells, and stamps certificates with structure + anchor coordinations;\nthree new theorem refs (<code>mkAnchoredFieldBcc</code>, <code>mkMeasuredFieldBcc</code>,\n<code>scaledAnchorsBccValid</code>) resolve-checked against the Lean sources.</li>\n<li>Flywheel wiring: <code>tools/mlip_local_promotion.py</code> loads the binding\nreport, re-checks every bound cell of the run&#39;s models through the Lean\nmirror, merges the certificates into the ODF gate&#39;s candidate metadata\n(they satisfy the formal-spec requirement without manual\n<code>--atlas-theorem-refs</code> flags), and carries a <code>field_certificates</code> block\nin the packet; <code>tools/mlip_phoenix_trace.py</code> emits it as\n<code>mlip.field_certificate</code> child spans (per-cell tier, exact scaled\nanchors, refusal witnesses, theorem ref) plus root-span rollups.\nNew tests: <code>tools/test_mlip_local_promotion.py</code>,\n<code>tools/test_mlip_phoenix_trace.py</code> (restores the justfile\n<code>flywheel-telemetry-check</code> target lost in the 2026-07-01 cleanup), both\nadded to the CI tools-smoke job.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> 64 bound cells (36 fcc + 28 bcc): 15 anchored softening\ninstances + 49 kernel-checked refusals. Of the bcc cells, 7 soften\nmonotonically end-to-end — chgnet on Cr/Fe/Mo/W, mace-mp-medium on Fe/Mo,\nand mace-mpa-0-medium on Nb — while all seven mace-mp-small bcc cells are\nrefused (anchors at or above bulk accuracy, the fcc noise-floor story\nrepeated). 181 build-locked theorems, 409 theorem declarations, 0 <code>sorry</code>.\nEnd-to-end smoke: a chgnet promotion packet now carries 16 per-cell\ncertificates (10 tier-2, 6 refusals) into its Phoenix spans, and the ODF\ngate promotes on certificate-backed formal fields alone.</li>\n<li><strong>Next.</strong> Extend the binder to the remaining Y-matrix structures (rocksalt\nMgO/NaCl and diamond Si need their own anchor layouts), and close the\nloop by reading refusal certificates back in the distill policy engine so\nrefused cells are excluded from correction at run time, not just flagged.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-10-measured-fields-rung-y-matrix-corpus-bound-into-per-cell-lean-field-instances-certificate-carrying-promotion-gate\">2026-07-10 - Measured fields rung: Y-matrix corpus bound into per-cell Lean field instances + certificate-carrying promotion gate</h2><ul>\n<li><strong>Why.</strong> The climate-series physics layer (same day, below) proved the\n<em>laws</em> of the environment error field abstractly; the discovery loop needs\nthose laws attached to <em>measured cells</em> so each (model, material) screening\noutput carries its own certificate, and needs refusals to reach runner\ntelemetry with their Lean witness instead of dying in a log line.</li>\n<li><strong>What.</strong><ul>\n<li><code>Theory/AnchoredField.lean</code>: the measurement bridge — clamped step\ninterpolation of the three fcc anchors (γ₁₀₀ → c=8, γ₁₁₁ → c=9,\nE_vac → c=11, bulk pin c=12) with generic monotonicity/softening proofs,\nso per-cell instances are one-liners with <code>norm_num</code> side goals; plus the\ndecidable <code>scaledAnchorsValid</code> admissibility predicate for kernel-checked\nrefusals.</li>\n<li>Two-tier field semantics, matching what the theorems actually need:\n<code>MeasuredField</code> (bulk pin only — closure, bulk-invariance, transfer, and\nnew measured-tier ranking-recovery laws) vs <code>ErrorField</code> (softening +\nmonotone — the directional barrier laws). Non-monotone cells (Ca/Sr\nanchor anomalies; MACE stiffening/noise-floor cells) keep certified\ncorrection semantics while the directional layer provably refuses them.</li>\n<li><code>python/scripts/bind_env_field_instances.py</code>: binds the 36 fcc Y-matrix\ncells (model-relaxed geometry + DFT-PBE targets) into per-atom anchors\n(Δγ · area-per-atom; ΔE_vac/12), emits\n<code>DistillAtlas/EnvFieldInstances.lean</code> (36 <code>MeasuredField</code> instances, 8\n<code>ErrorField</code> instances, 28 refusal theorems, all provenance-stamped,\ncorpus sha256 1f244b71846b) and a telemetry report\n(<code>data/y_matrix_runs/env_field_binding_report.json</code>).</li>\n<li><code>python/lupine_distill/odf/field_certificates.py</code>: the promotion-gate\nsurface — domain-gate admits/refusals with witness atoms, anchor\nadmissibility tiers, ranking-inversion impossibility certificates, each\nnaming its lean-spec theorem; <code>merge_into_candidate_metadata</code> feeds them\ninto <code>promotion_gate.evaluate_promotion</code> so decisions and telemetry carry\nmachine-checkable provenance. Tests\n(<code>python/tests/test_field_certificates.py</code>, 17 passing) pin the Python\npredicates to the same witness values the Lean modules lock and verify\nevery theorem reference resolves in the Lean sources.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green at 3661 jobs; 170 build-locked theorems,\n376 theorem declarations in <code>Materials/</code>, 0 <code>sorry</code>. Of 36 bound cells, 8\nexhibit monotone softening end-to-end (six chgnet cells — Ag, Al, Au, Cu,\nNi, Pd — plus Ni and Pt on mace-mp-medium); 28 carry kernel-checked tier-2\nrefusals — consistent with the proof pack&#39;s own account of noise-floor\nMACE-MPA-0 cells. Notable negative result: Ca and Sr break anchor\nmonotonicity for <em>every</em> model (DFT γ₁₁₁ ≈ γ₁₀₀ for these soft alkaline\nearths), a genuine limit of the smooth-decay idealization now recorded as\nrefusal certificates rather than silent misfits.</li>\n<li><strong>Next.</strong> Feed <code>field_certificates</code> outputs into the live flywheel\n(<code>tools/mlip_local_promotion.py</code> telemetry spans), and extend the binder\nbeyond fcc — the bcc anchor set ((100) → c=4, (110) → c=6, vacancy → c=7,\nbulk 8) needs its own <code>stepField</code> layout and admissibility predicate.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-10-climate-series-physics-layer-73-new-lean-theorems-for-the-five-material-classes\">2026-07-10 - Climate-series physics layer: 73 new Lean theorems for the five material classes</h2><ul>\n<li><strong>Why.</strong> The climate partnerships proof pack (&quot;The 0.2% Synthesis Problem&quot;,\n&quot;A Field, Not a Neural Net&quot;, &quot;Five Materials That Could Unlock 5–12\nGtCO₂/Year&quot;) commits the platform to five defect-mediated material classes —\ncobalt-free LMR cathodes, halide solid electrolytes, MOF DAC sorbents,\nelectrochemical ammonia catalysts, lead-free tin perovskites — but the\nformal evidence plane only covered fcc-metal elastic/error-geometry claims.\nThe underlying mechanics the pack argues from (softening ⇒ barrier\nunderestimation ⇒ exponential rate error ⇒ ranking inversion, plus the\ncorrection/impossibility dichotomy) existed as prose, not as machine-checked\nphysical law.</li>\n<li><strong>What.</strong> Added a seven-module physics layer to <code>lean-spec</code>\n(<code>Theory/EnvironmentField</code>, <code>Theory/BarrierArrhenius</code>,\n<code>Theory/RankingIntegrity</code>, <code>Theory/ScalingVolcano</code>,\n<code>Theory/DefectStability</code>, <code>Theory/SorptionStability</code>,\n<code>Validation/ClimateSeries</code>), 73 new theorems, all wired into the\n<code>Vision.lean</code> build locks (T78–T150) with six new <code>#guard</code> contracts.\nHighlights: the environment error field&#39;s seven structural laws (softening,\nbulk invariance, closure, family transfer, dominance, boundedness,\nzero-parameter blind continuation); the mechanism theorem\n(under-coordinated transition states ⇒ provably underestimated barriers);\nexact Arrhenius amplification identities with kernel-checked\nroom-temperature brackets (100 meV ⇒ 32–64× rate error; 180 meV ⇒ &gt;10³×\nconductivity error, riding Mathlib&#39;s log 2 decimal bounds); the\nmonotonicity impossibility lemma with a concrete cathode inversion witness;\nSabatier-volcano laws including non-monotonicity of activity in the\ndescriptor and breaker-flag soundness; the decidable metastability window\nhull-only screening provably misses; competitive-Langmuir humidity laws;\nand the first-shell domain gate with witnessed refusals.</li>\n<li><strong>Results.</strong> <code>lake build</code> green at 3659 jobs, 150 build-locked theorems,\n257 theorem declarations in <code>Materials/</code>, 0 <code>sorry</code>, 0 new axioms. The\nproof pack&#39;s quantitative claims (0.2% funnel, A-Lab novelty collapse, the\nkernel-rejected 27→26 episode, blind-prediction improvements, portfolio\nenvelope) are now decidable certificates in\n<code>Validation/ClimateSeries.lean</code>.</li>\n<li><strong>Next.</strong> Bind measured per-cell fields P(c) from the Y-matrix runs as\n<code>ErrorField</code> instances (the generated-evidence rung L2), so\n<code>corrected_recovers_reference_order</code> and the barrier theorems certify each\n(model, material) cell&#39;s screening output; then surface the domain-gate and\nimpossibility certificates through <code>promotion_gate.py</code> so refusals carry\ntheir Lean witness into runner telemetry.</li>\n</ul>\n<hr>\n<h2 id=\"2026-07-07-evidence-index-full-corpus-coverage-measured-retrieval-one-embedding-space\">2026-07-07 - Evidence index: full corpus coverage, measured retrieval, one embedding space</h2><ul>\n<li><strong>Why.</strong> CocoIndex was installed but under-used: it indexed only a handful of\nledger records, its semantic fast path never fired, nothing measured whether\nretrieval actually worked, and the largest prose corpora (research docs, the\npublished Library) were unindexed. We wanted the evidence tier to be\n<em>load-bearing</em> for the research loop, not just present.</li>\n<li><strong>What.</strong><ul>\n<li><strong>Coverage.</strong> Indexed the research-doc corpus (<code>docs/**</code>, root <code>*.md</code>), the\npublished Library sourced from the <em>deployed</em> <code>lupine-ledger</code> repo (with a\npending-vs-deployed publication-drift record), live-site <code>llms.txt</code> guides,\nand agent-run traces pulled back from Phoenix telemetry — one index spanning\nwhat agents did, what the program believes, and what we published.</li>\n<li><strong>Retrieval, measured.</strong> Built a 15-query gold-set eval (<code>eval_retrieval.py</code>).\nA/B&#39;d embedders and adopted <code>bge-small-en-v1.5</code> (384-dim); replaced SQL-LIKE\nkeyword search with FTS5/BM25; added hybrid (BM25+vector RRF) as the default\nmode. Activated the vec0 KNN sidecar that existed in code but was never built.</li>\n<li><strong>One space.</strong> Aligned the live GCP <code>evidence-index</code> service to the same\nbge-small model so the offline and live tiers share one embedding space.</li>\n<li><strong>Durability.</strong> Unit tests + CI on every <code>cocoindex/**</code> change; a nightly\nworkflow that refreshes all sources, rebuilds, and publishes retrieval\nmetrics + publication drift as tracked signals.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> Index grew 6 → ~2,880 chunks; no-change rerun stays 0.2s.\nRetrieval MRR: keyword 0.11 → 0.64 (BM25), semantic 0.45 → 0.51 (bge), best\nmode (hybrid+kind) <strong>0.82 / 0.93 hit@5</strong>. The eval caught two real defects\n(a rare-kind filter bug; the index claiming 5 unpublished articles as public)\nbefore any consumer hit them. vec0 KNN 1.4ms vs 20.9ms brute-force.</li>\n<li><strong>Next.</strong> Highest value: agents consult the index <em>before</em> acting\n(retrieval-augmented hypothesis generation) — the Literaturist is already a\nretrieval-augmented DO, so the pattern extends to Theorist/Causal. Then a\nclaim-lifecycle exporter (needs an <code>EvidenceChunk</code> metadata column so\n<code>--kind claim --status refuted</code> is a query). Edge-tier Vectorize memory is\nspecced and deferred (<code>docs/rfc-evidence-edge-memory.md</code>); the GCP redeploy to\nbge-small requires the re-embed backfill in <code>gcp/evidence-index/DEPLOY.md</code>.</li>\n</ul>\n<h2 id=\"2026-06-19-lupi-0-3-studio-molecule-trust-and-public-surface-split-prep\">2026-06-19 - LUPI 0.3 Studio, molecule trust, and public-surface split prep</h2><ul>\n<li><strong>Why.</strong> The checkout had a full LUPI release pass, public-surface split plan,\nLibrary export bundle, and molecule reliability work sitting locally. For a\nmajor release, that work needed to become one reviewable checkpoint instead of\nuntracked/in-progress workspace state.</li>\n<li><strong>What.</strong><ul>\n<li>Consolidated the LUPI viewer as <code>atlas-view@0.3.0</code>: mobile-first controls,\nlarger touch targets, TanStack Query saved-view caching, improved social\nshare metadata, picnic/cinematic sharing, first-class Lupi Studio 360 world\nbackgrounds, background grading controls, and optimized environment media.</li>\n<li>Kept MCP commands text/agent driven for this release; experimental voice\ncontrol is not part of the LUPI 0.3 surface.</li>\n<li>Added source-backed gallery nomenclature for PubChem-derived molecules,\nincluding PubChem CIDs, formulas, systematic names, aliases, and a local\nreliability/backup audit tool.</li>\n<li>Added the public-surface repo-split map and extraction packets for\n<code>lupine.science</code>, <code>lupi.live</code>, <code>library.lupine.science</code>, and the remaining\nscience/control-plane repo.</li>\n<li>Added <code>scripts/export_library_content.mjs</code> and generated the first\n<code>exports/library-content/latest</code> bundle for the Library extraction path.</li>\n<li>Recorded <code>exports/</code> in the root ownership ledger so generated public\nartifacts have an explicit owner.</li>\n</ul>\n</li>\n<li><strong>Results.</strong><ul>\n<li><code>pnpm audit:nomenclature</code>: 55 gallery entries, 24 nomenclature records, 0\nerrors, 8 non-blocking provenance warnings for older procedural/synthetic\nexamples.</li>\n<li><code>pnpm --filter @atlas/ui test -- src/backgroundPresets.test.ts src/store.test.ts src/gallery-data.test.ts</code>: 44 tests passed.</li>\n<li><code>pnpm --filter @atlas/ui build</code>: passed.</li>\n<li><code>pnpm --filter @atlas/web build</code>: passed and generated 8 static SEO routes.</li>\n<li><code>pnpm verify:gallery --no-screenshot</code>: 20/20 checks passed, including real\ndataset load.</li>\n<li><code>pnpm verify:controls --no-screenshot</code>: desktop controls smoke passed.</li>\n<li><code>pnpm verify:controls:mobile --no-screenshot</code>: mobile controls smoke\npassed.</li>\n<li><code>VERIFY_URL=http://127.0.0.1:5174/#/mcp pnpm verify:mcp-bridge</code>: passed\nagainst the built web output served locally.</li>\n<li><code>npm --prefix library-site run build</code>: built 55 articles.</li>\n</ul>\n</li>\n<li><strong>Next.</strong><ul>\n<li>Open the release PR and keep local/CI/deploy/live truth separate.</li>\n<li>Resolve the eight older gallery provenance warnings before treating the\nentire curated gallery as source-backed.</li>\n<li>Promote the public-surface split only after each extracted repo has its own\nCI, deploy, secrets, and live health proof.</li>\n</ul>\n</li>\n</ul>\n<h2 id=\"2026-06-16-academic-review-of-the-projection-law-immi-suite-first-fix-pass\">2026-06-16 — Academic review of the Projection Law / IMMI suite, first fix pass</h2><ul>\n<li><strong>Why.</strong> An independent adversarial review flagged six MUST-FIX gates before any journal submission: the affine/smooth/finite-sample theorems were being oversold as deriving global claims; the MLIP factorial result needed permutation-floor nuance; theorem counts and PDF URLs were inconsistent across surfaces; and ORCID/DOI placeholders were still open.</li>\n<li><strong>What.</strong><ul>\n<li>Confirmed the formal core is unchanged and green: <code>lean-spec lake build</code> (2891 jobs, 77 build-locked theorems in <code>Vision.lean</code>, 0 <code>sorry</code>, 0 new axioms).</li>\n<li>Tightened the PRX and IMMI manuscripts: qualified the affine/gauge and smooth-local/global bridges, downplayed the finite-sample PR sample-complexity claim, and added permutation-floor wording to the abstract, Table 1, and the IMMI abstract.</li>\n<li>Rebuilt both PDFs (<code>paper2/projection-law.pdf</code>, <code>paper2/immi/projection-law-immi.pdf</code>) and submission bundles.</li>\n<li>Published the academic review at <code>docs/reviews/academic-review-projection-law-2026-06-16.md</code> and surfaced it as a first-class article on <code>library.lupine.science</code> via <code>library-site/scripts/catalog.js</code>.</li>\n<li>Added versioned PDF assets under <code>library-site/src/assets/papers/</code> and updated <code>working-papers.html</code> to serve them from the site instead of a stale GCS URL.</li>\n<li>Prepared submission collateral: <code>paper2/cover-letter.md</code>, <code>replication/error-geometry/.zenodo.json</code>, and updated <code>ZENODO_DEPOSIT.md</code>.</li>\n</ul>\n</li>\n<li><strong>Results.</strong><ul>\n<li><code>paper2/python quality_gate.py --lean</code> passes.</li>\n<li><code>git diff --check</code> clean.</li>\n<li>Library site builds 54 articles; the review renders at <code>/#/read/academic-review-projection-law</code>.</li>\n</ul>\n</li>\n<li><strong>Next.</strong><ul>\n<li>Fill the author ORCID and mint the Zenodo DOI (user action).</li>\n<li>Run the adversarial multi-agent review pass described in <code>TARGETING.md</code> and incorporate any findings.</li>\n<li>Update the external <code>lupine.science</code> marketing page (source not in this repo) to point to the new versioned PDF.</li>\n</ul>\n</li>\n</ul>\n<h2 id=\"2026-06-14-lupi-controls-palette-rollout\">2026-06-14 - LUPI controls palette rollout</h2><ul>\n<li><strong>Why.</strong> The viewer&#39;s advanced visual controls had outgrown the fixed side drawer. Users needed\na more deliberate inspection surface: one place to tune look, surface, world, and export settings\nwithout covering the molecule or forcing repeated drawer navigation.</li>\n<li><strong>What.</strong><ul>\n<li>Replaced the desktop controls drawer with a dockable, resizable, collapsible tool palette.</li>\n<li>Consolidated Look, Surface, World, and Export into one tabbed Controls surface.</li>\n<li>Kept the mobile bottom sheet path intact while making desktop panel chrome consistent.</li>\n<li>Added a portless browser verification harness, <code>pnpm verify:controls</code>, that starts Vite on a\nfree local port, loads a real C60 molecule, checks all four control tabs, validates there is\nonly one close affordance per embedded panel, and exercises resize plus collapse/expand.</li>\n</ul>\n</li>\n<li><strong>Results.</strong><ul>\n<li>Local production build is green: <code>pnpm --filter @atlas/web build</code>.</li>\n<li>Portless controls smoke is green on random local ports: <code>pnpm verify:controls -- --no-screenshot</code>.</li>\n<li>The rollout no longer depends on a manually running fixed-port dev server for verification.</li>\n</ul>\n</li>\n<li><strong>Next.</strong><ul>\n<li>Ship through the normal push-to-main Cloud Run viewer deploy and verify <code>https://lupi.live</code>\nagainst the same controls harness via <code>VERIFY_URL</code>.</li>\n<li>Add screenshot diffing once the visual-regression lane graduates from artifact capture to\nassertions.</li>\n</ul>\n</li>\n</ul>\n<h2 id=\"2026-06-12-repo-consolidation-and-onboarding-sprint\">2026-06-12 — Repo consolidation and onboarding sprint</h2><ul>\n<li><strong>Why.</strong> The repo had grown three overlapping Distill roots (<code>distiller/</code>, <code>lupine-distill/</code>, <code>atlas-distill/</code>) and a maze of docs that made it hard for new developers and research scientists to find the right entry point. The one-root-one-owner rule was being violated, and stale paths were still referenced by active code.</li>\n<li><strong>What.</strong><ul>\n<li>Consolidated Distill by runtime/language: <code>atlas-distill/</code> is the single Rust engine, <code>python/</code> is the single home for active Python packages, and retired material moved to <code>archive/</code> (<code>distiller-kb/</code>, <code>lupine-distill-rust/</code>, <code>lupine-dspy/</code>, <code>tools-retired/</code>).</li>\n<li>Updated every active import path and sys.path hack from the old roots to <code>python/</code>.</li>\n<li>Added onboarding docs: <code>docs/ONBOARDING.md</code>, <code>docs/ARCHITECTURE.md</code>, <code>docs/GLOSSARY.md</code>, <code>docs/FAQ.md</code>, <code>CONTRIBUTING.md</code>, <code>docs/HANDBOOK.md</code>, and per-root READMEs for <code>python/</code>, <code>gcp/</code>, <code>data/</code>, <code>atlas/</code>, <code>mlip_immi/</code>, <code>lean-spec/</code>, <code>glim-think/</code>, <code>library-site/</code>, <code>paper/</code>, and <code>tools/</code>.</li>\n<li>Added <code>scripts/bootstrap.ps1</code> / <code>bootstrap.sh</code>, extended the <code>justfile</code> with <code>verify</code>, <code>bootstrap</code>, <code>bootstrap-heavy</code>, and <code>docs-serve</code> recipes, and created GitHub issue/PR templates.</li>\n<li>Added <code>.github/workflows/verify.yml</code> (Python unit tests, Rust check, tools smoke, diff hygiene) and hardened <code>python/pyproject.toml</code> with realistic dependency extras matching the MLIP backend images.</li>\n<li>Audited <code>docs/</code> for stale/provisional files and added honest banners with redirects to current sources.</li>\n</ul>\n</li>\n<li><strong>Results.</strong><ul>\n<li><code>just verify</code> passes locally: Python unit tests (92 passed), Rust check, tools smoke tests (18 passed, 4 skipped), and <code>git diff --check</code> clean.</li>\n<li><code>cargo test --manifest-path atlas-distill/Cargo.toml --bin atlas-distill</code> passes (100 tests).</li>\n<li>No hardcoded secrets found in active code; env vars are used for all tokens/keys.</li>\n</ul>\n</li>\n<li><strong>Next.</strong><ul>\n<li>Continue migrating any remaining active <code>lupine-distill</code> / <code>distiller</code> references in historical docs.</li>\n<li>Add a Lean-proof CI job (cached) once the <code>lean-spec</code> first-build cost is acceptable.</li>\n<li>Consider moving reusable <code>mlip_immi/</code> logic into <code>python/lupine_distill/</code> with tests.</li>\n</ul>\n</li>\n</ul>\n<h2 id=\"2026-06-02-correction-retracting-bounding-the-day-39-s-ribbon-overclaims\">2026-06-02 — CORRECTION: retracting / bounding the day&#39;s ribbon overclaims</h2><ul>\n<li><strong>Why.</strong> A working session produced several confident claims that do <strong>not</strong> hold up.\nLogging the retraction here because self-correction is the method, and the entries below\n(and a now-removed framing) overstated results — mostly because the work was done against a\n<strong>stale CSV export and the MPtrj MLIP-energy lane, disconnected from the live D1 ledger\ncorpus</strong> the program actually runs on (OpenKIM/NIST elastic constants). Orientation to the\nlive system is now captured in the <code>lupine-system-architecture</code> skill + <code>docs/science/objects.md</code>.</li>\n<li><strong>Retracted / bounded:</strong><ol>\n<li><em>&quot;Generalised the ribbon, first-principles&quot; (<code>RibbonProjection.lean</code>)</em> — <strong>withdrawn.</strong> That\nmodule is a <strong>generic scalar operative-value lemma</strong> (the same parabola already in\n<code>ContextSpecificProof</code>), <strong>not</strong> a formalization of the hyper-ribbon (model manifold), the\nparticipation-ratio measure, or the keystone configuration-space core. Its docstring is\ncorrected to say so; it does not &quot;generalise the ribbon.&quot;</li>\n<li><em>&quot;<code>broad_commitment_is_open</code> is TRUE under per-backend policy selection&quot;</em> — <strong>bounded to\nnear-meaningless.</strong> That campaign measured <strong>energy-MAE on MPtrj DFT rows</strong>, a different lane\nfrom the OpenKIM/NIST elastic-constant corpus the ribbon is built on; the distill &quot;win&quot; is an\n<strong>energy-block recalibration that leaves forces/stress/elastic unchanged</strong>, so it does not\nimprove the potential for any force-driven use. Not a model improvement. See the banner on\n<code>docs/glim-m3-upgrade/runs/live-campaign-results.md</code>.</li>\n<li><em>&quot;A6 has real but conditional support&quot;</em> — <strong>demoted to untrusted.</strong> The A6 test was a\n5-structure MLIP-force pilot on the wrong lane <strong>without the mandatory coupling-aware null</strong>\n(Cauchy relation / mechanical stability — Jackson–Somers / Archie). Suggestive at most; must\nbe redone on the live corpus with the coupling-aware null before it is believed.</li>\n<li><em>&quot;category error, verified in code&quot;</em> — <strong>overstated</strong> (corrected in the banner atop\n<code>docs/science/keystone-reconciliation.md</code>): the repo&#39;s PR-of-error-covariance is the standard\nsloppy-model effective-dimensionality measure, not an elementary mistake.</li>\n</ol>\n</li>\n<li><strong>What still stands:</strong> the MiniMax M2.7→M3 <strong>model-axis</strong> engineering (typechecked, tested); the\n<strong>documentation architecture</strong> (<code>docs/navigation.md</code>, <code>docs/science/objects.md</code>, ADR-0002); and\nthe reusable analysis <strong>tools</strong> (they were just pointed at the wrong data).</li>\n<li><strong>Next.</strong> Redo Q1/Q2/Q3 on the <strong>live D1 ledger + GCS lake</strong> (not exports), per-element /\nmatched-n, with coupling-aware nulls, writing results back into the ledger.</li>\n</ul>\n<h2 id=\"2026-06-02-back-to-the-keystone-paper-the-category-error-and-the-first-test-of-a6\">2026-06-02 — Back to the keystone paper: the category error, and the first test of A6</h2><ul>\n<li><strong>Why.</strong> The &quot;promote per-backend distill&quot; framing was disinterested box-ticking — it ignored\nthat forces never moved. Went back and actually read <em>A Conditional Universality Theorem for\nError Geometry in MLIPs</em> (repo-root PDF) to ground the program in its own theory.</li>\n<li><strong>What.</strong> A reconciliation of the repo&#39;s ribbon claims against the paper\n(<code>docs/science/keystone-reconciliation.md</code>), plus the <strong>first direct empirical test of the paper&#39;s\nload-bearing assumption A6</strong> (&quot;common-spatial-mode separability&quot;) — <code>tools/a6_alignment_test.py</code>,\nrun on the force-error field (3 MLIPs × 5 shared structures = 107 atoms, 5000 stratified\npermutations), with three statistics vs a within-structure permutation null.</li>\n<li><strong>Results.</strong> Two findings. (1) <strong>Category error, verified in code.</strong> The repo computes the\nparticipation ratio of the error-<em>vector</em> covariance in <em>observable</em> space\n(<code>manifold.rs:43-112</code>, 3×3 over C11/C12/C44) and calls it a low-dimensional &quot;error manifold.&quot; The\npaper&#39;s theorem is about a <em>configuration-space</em> core <code>H ⊂ ℝᵐ</code>; its error <em>boundary</em> is dimension\n<code>m−1</code>, <strong>not</strong> low; and a measure-theoretic concentration must not be called a manifold. The\nbridge between the two (A6) was <strong>assumed, never stated</strong> — including in my own\n<code>RibbonProjection.lean</code>, which formalizes the wrong (toy) object. (2) <strong>A6 has real but\nconditional support.</strong> All three MLIPs concentrate force error on the <strong>same atoms</strong>\n(<code>mag_corr</code> 0.70–0.86, p≤0.0002, well above the stratified null ~0.34) and err in correlated\ndirections (<code>atom_cos</code> 0.2–0.3, sig) — first force-level evidence for shared structure. But it&#39;s\nheterogeneous: <strong>CHGNet is a partial outlier</strong> (whole-field alignment with MACE n.s., p=0.09),\nindependently echoing that CHGNet is the backend distill regressed. So it&#39;s the paper&#39;s\n<em>perturbative/conditional</em> regime, not the unrestricted ribbon claim.</li>\n<li><strong>Next.</strong> Scale the A6 test to MatPES/MPtrj/OMat24 with a blocked bootstrap over materials (the\npaper&#39;s protocol); estimate per-model perturbation <code>δ_M</code> (CHGNet largest); and formalize the\npaper&#39;s actual <code>exact_tubular_universality</code> skeleton (reach + 1-D monotonicity) instead of the\n<code>RibbonProjection</code> toy.</li>\n</ul>\n<h2 id=\"2026-06-02-live-3-tier-sim-campaign-distill-is-energy-only-per-backend-policy-gated\">2026-06-02 — Live 3-tier sim campaign: distill is energy-only + per-backend-policy-gated</h2><ul>\n<li><strong>Why.</strong> The M3 upgrade work configured a 3-tier Cloud-Run sim matrix (baseline /\ndistill-accuracy / distill-accuracy+speed) but had not <em>run</em> it. Run it for real on\nGPUs to test the local-Opus T4 hypotheses about where the ribbon distill correction\nfails, and to put hard numbers on <code>broad_commitment_is_open</code>.</li>\n<li><strong>What.</strong> 23 cells on <code>shed-489901</code> L4 jobs (<code>mlip-cell-{mace,chgnet,sevennet}</code>),\n3 tiers × 3 backends × {energy_volume, forces}, plus global-tuned and per-backend-tuned\nre-runs. Scored by <code>tools/mlip_sim_matrix.py</code> (Lean <code>cellValue</code>). Surfaced + fixed a real\nbug (runner wants catalog id <code>mace-mp-0</code>, not <code>mace</code>). Cost ≈ $2. Full table:\n<code>docs/glim-m3-upgrade/runs/live-campaign-results.md</code>.</li>\n<li><strong>Results.</strong> (1) <strong>Baselines reproduce exactly</strong> (MACE 0.41161, SevenNet 0.3997) and the\ntuned MACE cell returns <strong>0.2038</strong> — the committed <code>maceEnergyDistill</code> to 4 decimals, same\npolicy hash. The harness is faithful. (2) <strong>Distill is energy-only:</strong> on <code>forces</code> the\ncorrection changes the error by <strong>0.0%</strong> for all three backends — an exact confirmation of\nthe pre-registered local-Opus hypothesis <strong>T4-H2</strong> (energy/mechanical orthogonality).\n(3) <strong>Distill is policy-gated, not automatic:</strong> a generic policy regresses the\nalready-accurate CHGNet (0.1035 → 0.1429, −38%); the global tuned policy still regresses it\n(−28%); but CHGNet&#39;s <strong>own</strong> <code>signed-orientation</code> policy flips it to <strong>+6.1%</strong> (0.0971).\nNet: distill beats baseline on all three backends <strong>iff each uses its own policy</strong> — TRUE\nunder per-backend selection, FALSE under any single global policy.</li>\n<li><strong>Next / done same session.</strong> <strong>Generalised the ribbon, first-principles</strong> —\n<code>lean-spec/.../Theory/RibbonProjection.lean</code>. Rather than encode the campaign as per-backend\ncases (the exception handling we explicitly avoid), it proves all three findings as corollaries\nof one model-independent geometry: error = ribbon-parallel <code>par</code> ⟂ orthogonal <code>orth</code>, a scalar\ncorrection <code>κ</code> acts only on <code>par</code>, and <code>ribbonGain = κ·(2·par − κ)</code> (the orthogonal sector\ncancels). Corollaries: <code>orthogonal_error_gain_nonpos</code> (energy-only / forces 0%),\n<code>ribbonGain_neg_of_antialigned</code> (CHGNet regression = misaligned κ, same parabola — no model\naxiom), <code>ribbonGain_strictly_valuable</code> (MACE/SevenNet), and the capstone\n<code>broad_value_no_model_exception</code> (two backends, same <code>par</code>, aligned κ ⇒ equal positive gain).\nAll proofs are <code>ring</code>/<code>nlinarith</code> over arbitrary reals; algebraic identities independently\nverified in sympy (all green) <strong>and the module is kernel-verified locally</strong> — <code>lake build</code>\ngreen, 0 sorry, after fetching the Mathlib cache. A second live result sharpened &quot;energy-only&quot;\nto <strong>&quot;support-set-only&quot;</strong>: a stress-targeted policy also left stress ≈ unchanged, and the\n<code>elastic_constants</code> distill cell <em>refused</em> (<code>requires ≥6 cases; found 0</code>) — the correction can\nonly move a property the support manifold covers, exactly <code>RibbonProjection</code>&#39;s\northogonal-sector law. Open: accelerate speed axis on <code>elastic_constants</code> with the\nelastic-covering <code>train-plus-elastic-v1</code> support (re-running).</li>\n</ul>\n<h2 id=\"2026-06-02-theorist-deep-tier-model-upgrade-minimax-m2-7-m3-gated-on-a-measured-a-b\">2026-06-02 — Theorist deep-tier model upgrade: MiniMax M2.7 → M3, gated on a measured A/B</h2><ul>\n<li><strong>Why.</strong> MiniMax-M3 (released 2026-06-01) is the new top-tier deep model behind Theorist\nhypothesis generation — 1M context, ~1/20 cost at long context, same <code>api.minimax.io/v1</code>\nroute. But a model swap is only trustworthy if the quality lift is <em>measured</em> against a fixed\nresearch target, not assumed. The worker could pin a provider but not a specific MiniMax model,\nso M2.7-vs-M3 was not even expressible.</li>\n<li><strong>What.</strong> (1) Added a <strong>model axis</strong>: <code>selectDeepRoute({modelOverride})</code> →\n<code>generateResearchText({modelOverride})</code> → <code>miniMaxModel(env, id)</code>; <code>/ops/experiment-generate</code>\naccepts <code>body.model</code>; <code>ab-oracle.ts --axis model</code>. Default flips to <code>MiniMax-M3</code> with\n<code>MINIMAX_BASELINE_MODEL = &quot;MiniMax-M2.7&quot;</code> kept as the canonical A/B baseline (rollback = a\n<code>MINIMAX_MODEL</code> secret change, no redeploy). (2) Pinned the <strong>ribbon target theorem set</strong>\n(T1 <code>hyper_ribbon_bound_3d</code>, T2 <code>empirical_hyper_ribbon_holds</code>, T3 <code>ParameterBound</code>, T4\n<code>broad_commitment_is_open</code>, the <code>cellValue</code> bridge) and a per-theorem <strong>research strategy</strong>.\n(3) Built the <strong>eval harness</strong>: <code>glim-ribbon-theorems</code> dataset + <code>tools/glim_model_eval.py</code>\n(generate→compare→report, mirrors ab-oracle&#39;s adopt/reject) with a <strong>local-Opus</strong> rubric.\n(4) Configured the <strong>3-tier Cloud-Run sim matrix</strong> (<code>policies/model-sim-matrix.yml</code> +\n<code>tools/mlip_sim_matrix.py</code>) — baseline / distill-accuracy / distill-accuracy+speed, cost-bounded\nand <code>cellValue</code>-scored. Full process in <code>docs/glim-m3-upgrade/</code>.</li>\n<li><strong>Results.</strong> Model-axis change typechecks with <strong>0 new type errors</strong> (proven against the committed\noriginal) and existing tests stay green (6/6); repo <code>lint:fast</code> clean. The <strong>local Opus agent</strong>,\nrun for real, generated rigorous competing hypotheses for T1/T3/T4 and — judging blind — scored\nthem 10/10 against a deliberately weak fixture at 0/10 (rubric discriminates). The sim driver,\nscored on the committed MACE-energy artifacts, returns <strong>cellValue 1.238</strong> for distill_accuracy\n(50.5% MAE cut, speedup 1.025) and reproduces the <code>AccuracyCommitment.lean</code> constants from the raw\nartifact. No M2.7/M3 generation numbers were fabricated (this checkout has no MiniMax key).</li>\n<li><strong>Next.</strong> (1) With <code>INTERNAL_TASK_TOKEN</code> + a live key, run the M2.7→M3 generation and let the same\nlocal-Opus judge emit the real verdict. (2) Launch the canary sim matrix on <code>Cu_fcc</code>/<code>Si_diamond</code>/\n<code>Fe_bcc</code> to test the T4 &quot;where does distill fail?&quot; hypotheses and bound <code>broad_commitment_is_open</code>.</li>\n</ul>\n<h2 id=\"2026-05-29-neural-symbolic-loop-gpu-mlip-curvature-machine-checked-lean-0-sorry\">2026-05-29 — Neural-symbolic loop: GPU MLIP curvature → machine-checked Lean (0 sorry)</h2><ul>\n<li><strong>Why.</strong> Close the proof↔physics gap at the tightest coupling — a number measured on the GPU\nbecoming a theorem the Lean kernel checks the next moment — and seed <code>atlas_theorems</code> <em>from the\nphysics</em>, not by hand.</li>\n<li><strong>What.</strong> A three-node continuous loop (<code>python/scripts/neural_symbolic/</code>):\n<strong>Node 1</strong> pits MACE-MP-0 vs CHGNet on a pure-shear C44 strain sweep of FCC Ni (the curvature\nobservable) on the A4500; <strong>Node 2</strong> relays T3-REJECT breaches as OpenInference spans (the Python\nflywheel pattern — entirely off the glim-think <code>tsc</code> path; live OTLP when <code>PHOENIX_OTLP_RELAY_URL</code>\nis set, durable local artifact otherwise); <strong>Node 3</strong> authors import-free core-Lean theorems\n(<code>by decide</code>, 0 sorry) encoding the empirical failure as a verified negative constraint, plus an\n<code>atlas_theorems</code> seed.</li>\n<li><strong>Results.</strong> On the GPU: MACE-MP-0 elastic C44 <strong>92.4 GPa (−25.9%) → REJECT</strong>, CHGNet <strong>101.2 GPa\n(−18.8%) → REVIEW</strong> vs the 124.7 GPa literature reference (MACE units cross-validated against the\ntorch_sim elastic run). The MACE breach was authored into\n<code>mace_mp_0_curvature_reject : 323*4 &gt; 1247 := by decide</code> and <strong>independently verified by a fresh\n<code>lean</code> compile (rc 0, 0 sorry)</strong>. 4 theorems synthesized, 4 <code>atlas_theorems</code> seed rows\n(<code>status=&#39;verified&#39;</code>). Report: <code>docs/neural-symbolic-curvature-loop.md</code>.</li>\n<li><strong>Next.</strong> (1) Stand up the GCP OTLP relay for live Phoenix streaming (then the loop is fully\ncontinuous: GPU → Phoenix → Lean per measurement). (2) Widen Node 1 to the full phonon Hessian.\n(3) Apply the seed to the live <code>glim-ledger</code> D1.</li>\n</ul>\n<h2 id=\"2026-05-29-local-gpu-proof-torchsim-distill-uplift-formal-gate-ni-fcc\">2026-05-29 — Local GPU proof: TorchSim → distill → uplift → formal gate (Ni FCC)</h2><ul>\n<li><strong>Why.</strong> The ATLAS PR wired the rails but ran no train: Track B&#39;s TorchSim backend was a\nstub, the formal promotion gate had never scored a real benchmark, and the cloud &quot;~75% Ni\nzero-point ribbon lift&quot; was unreproduced locally. With a real GPU (RTX A4500) on hand, prove\nthe whole compute loop end to end.</li>\n<li><strong>What.</strong> Stood up a CUDA env (torch 2.6.0+cu124, torch_sim 0.6.0, cached MACE-MP-0) and wrote\n<code>python/scripts/run_ni_gpu_loop.py</code> — the GPU runner the Track B stub\ndefers to. It benchmarks MACE-MP-0 on the sealed Ni FCC EAM fixture via TorchSim, fits the\nzero-point distill correction on the <em>non-overlapping</em> support set, computes real elastic\nconstants + <code>distill_v_uplift</code>, and drives the ATLAS formal gate. Also filled\n<code>TorchSimBenchmarkBackend.run()</code> for real (70 Track-B tests still green; CI-safe without torch_sim).</li>\n<li><strong>Results.</strong> 31 eval structures in 5.6 s on the A4500. Energy MAE vs Mishin EAM <strong>1.2803 →\n0.0037 eV/atom (99.7%)</strong> after the +1.279 eV/atom zero-point correction; stress 0.861 → 0.274 GPa;\n<strong>overall <code>distill_v_uplift</code> 76.0%</strong> — independently reproducing the cloud material-family\nresult. Real elastic constants: MACE-MP-0 <strong>C11=262.9 (+6.7%), C12=166.6 (+13%), C44=92.4\n(−26%)</strong> GPa vs literature (the EAM-reference recovery returned 246.5/147.3/124.7 exactly,\nvalidating the fit). Formal gate, full range on real numbers: in-support certified →\n<strong>PROMOTE</strong>; out-of-support (the proved T3 negative-transfer regime) → <strong>REVIEW</strong>;\nmarginal+uncertified → <strong>REJECT</strong>. The formal layer demonstrably gates real GPU compute.\nReport: <code>docs/mlip-gpu-ni-distill-formal-gate.md</code>.</li>\n<li><strong>Next.</strong> (1) Genuine cross-material negative transfer (second MLIP/material) for a real T3\nREJECT. (2) Curvature lane — the C44 shear undershoot points straight at the phonon/Hessian\nfrontier. (3) Seed glim-think <code>atlas_theorems</code> + emit <code>lupine.proof.status</code> spans so this gate\ndecision flows into Phoenix.</li>\n</ul>\n<h2 id=\"2026-05-29-atlas-lean-integration-formal-foundations-closed-loop-scaffolding\">2026-05-29 — ATLAS-Lean integration: formal foundations + closed-loop scaffolding</h2><ul>\n<li><strong>Why.</strong> Meta open-sourced ATLAS-Lean (autoformalized textbook mathematics) and <code>torch-sim</code>.\nThe <code>ATLAS_Lean_Integration_Review</code> laid out a 7-phase plan to (a) put <code>lean-spec</code> on a shared,\nreproducible Mathlib by pinning to ATLAS&#39;s revision, (b) make the MLIP benchmark/distill loop\nmeasurable, and (c) thread formal foundations through glim-think, Phoenix, the ODF promotion\ngate, and dspy.</li>\n<li><strong>What.</strong><ul>\n<li><em>lean-spec (Phase 1+2).</em> Pinned the toolchain to ATLAS&#39;s <code>v4.29.0</code> and Mathlib to <code>8a178386…</code>,\nadded <code>facebookresearch/atlas-lean</code> as a Lake dependency at <code>c5a10f1a</code>. Added 6 ATLAS-backed\ntheorems (Jacobian rank ≤ P / ≤ N / ≤ min(P,N) — the formal core of the Parameter-Bound\nconjecture — plus ℝ-level hyper-ribbon corollaries) to <code>Analysis.Manifold</code> and\n<code>Theory.ParameterBound</code>.</li>\n<li><em>lupine-distill (Track B).</em> <code>TorchSimBenchmarkBackend</code> (lazy import + Mock fallback), the\n8-benchmark suite, the canonical <code>BenchmarkResult</code>/<code>BenchmarkMetrics</code> schema, and the\n<code>distill_v_uplift</code> calculator with promote/review/reject gates; 70 tests.</li>\n<li><em>gcp/mlip-cell-runner (Track C).</em> Consolidated the per-backend &quot;lone wolf&quot; sprawl into one\n<code>pyproject.toml</code> + matrixed Dockerfile/cloudbuild, scheduled-run policies, and an OpenInference\npatcher + loop connector; legacy files retained with a <code>DECOMMISSION.md</code> map.</li>\n<li><em>glim-think (Track D).</em> <code>atlas_theorems</code> D1 migration (<code>0010</code>), per-facet ATLAS context loader,\nOpenInference span-kind + ATLAS-attribute telemetry helpers, and three Phoenix eval runners.</li>\n<li><em>distiller/ODF + lupine-dspy (Track E).</em> Formal-verification promotion gate + theorem-aware\nOperatorPack model card + schema bridge; <code>TheoremGuidedHypothesis</code> dspy signature +\nformal-provenance persistence migration; 40 tests.</li>\n</ul>\n</li>\n<li><strong>Results.</strong> <code>lake build</code> green (1502 jobs, <strong>96 theorems, zero <code>sorry</code></strong>) under the\naligned/downgraded Mathlib — every existing proof survived the pin. 110 Python tests pass across\nB/E (independently re-run). <strong>Key finding:</strong> ATLAS&#39;s autoformalized modules elaborate at ~7–9 min\n<em>each</em>; importing a whole subject (<code>Atlas.RealAnalysis</code>, ~85 modules) is ≈80 min and twice\nexhausted the dev machine&#39;s memory. ATLAS <em>does</em> compile cleanly in-workspace (71/85 modules built\nwith zero errors before a reset), so the dependency is wired and viable — but whole-subject imports\nare cost-prohibitive, so the new theorems build on the shared cached Mathlib that ATLAS pins to.</li>\n<li><strong>Next.</strong> (1) Selective ATLAS <em>leaf-module</em> imports behind an offline/opt-in build target so CI\nstays fast. (2) Apply <code>0010</code> to the LEDGER D1 and seed real theorem rows from <code>lean-spec</code>.\n(3) Wire <code>distill_v_uplift</code> into the ODF promotion gate on a real model pair. (4) Resolve the ORB\ncu118 / UMA numpy-2 stack split flagged in the GCP <code>DECOMMISSION.md</code> before deleting legacy reqs.</li>\n</ul>\n<h2 id=\"2026-05-18-fix-mislabeled-home-page-working-paper-banner\">2026-05-18 — Fix mislabeled home-page working-paper banner</h2><ul>\n<li><strong>Why.</strong> The Library home banner promoted the paper link as <em>&quot;Immigrant Scientist — The\nInvisible Foundation — a data-driven analysis of immigrant contributions to US science.&quot;</em>\nThe author is not an immigrant and that is not the paper. The copy was a confused\nmisreading of an internal IMMI working label as &quot;immigrant.&quot;</li>\n<li><strong>What.</strong> Verified <code>/immi_paper.pdf</code> is in fact <em>The Causal Geometry of Prediction Errors\nin Interatomic Potentials</em> (Welcing, Lupine Science). Corrected the\n<code>home.preprint.*</code> strings (EN + ZH) in <code>i18n.js</code> to the real title/abstract; the link was\nalways correct.</li>\n<li><strong>Results.</strong> Banner now identifies the paper as a working paper in preparation:\n<em>The Causal Geometry of Prediction Errors in Interatomic Potentials</em>. No &quot;immigrant&quot; copy\nremains in the build.</li>\n<li><strong>Next.</strong> Audit other recovered hardcoded copy for the same era of stale text.</li>\n</ul>\n<h2 id=\"2026-05-19-paper-build-auto-dispatches-the-library-deploy\">2026-05-19 — paper-build auto-dispatches the Library deploy</h2><ul>\n<li><strong>Why.</strong> First real run of <code>paper-build.yml</code> proved a gap: a push made with a\nworkflow&#39;s <code>GITHUB_TOKEN</code> does not trigger other workflows (GitHub&#39;s recursion\nguard), so the rebuilt PDF landed on <code>main</code> but <code>deploy-library-site</code> never fired —\nit needed a manual nudge.</li>\n<li><strong>What.</strong> Added <code>actions: write</code> permission and an explicit\n<code>gh workflow run deploy-library-site.yml --ref main</code> after the push (skipped on the\nbyte-identical no-op path).</li>\n<li><strong>Results.</strong> <code>gh workflow run paper-build.yml</code> is now genuinely one command:\ncompile → quality gate → commit → deploy. Verified the first run&#39;s PDF live (fresh\nrecompile from the newer <code>.tex</code>; abstract now carries the foundation-MLIP transfer\nresult; 0 raw-LaTeX leaks).</li>\n<li><strong>Next.</strong> None — the paper pipeline is closed-loop and opt-in.</li>\n</ul>\n<h2 id=\"2026-05-18-opt-in-ci-to-rebuild-the-working-paper-pdf\">2026-05-18 — Opt-in CI to rebuild the working-paper PDF</h2><ul>\n<li><strong>Why.</strong> The broken-PDF fix was a one-time swap; the root cause — a stale/broken\nlocal PDF can be the served artifact — remained. But the paper shouldn&#39;t rebuild on\nevery push; the author chooses when to upgrade it.</li>\n<li><strong>What.</strong> Added <code>.github/workflows/paper-build.yml</code>, <strong><code>workflow_dispatch</code> only</strong>\n(never on push). Run it with <code>gh workflow run paper-build.yml</code>. It compiles\n<code>paper/immi-paper.tex</code> with the committed figures (optional <code>regenerate_figures</code>\ninput), enforces a <strong>quality gate</strong> that fails the run if the PDF text layer leaks\nraw LaTeX (<code>\\textbf</code>, <code>\\noindent</code>, <code>$C_{</code>) or mojibake or is &lt; 5 pages — the exact\n<code>immi-paper-local</code> class of bug — and, when <code>deploy=true</code> (default), commits the\nfresh PDF to <code>main</code> so the existing Library deploy publishes it. <code>deploy=false</code>\nbuilds an inspectable artifact only.</li>\n<li><strong>Results.</strong> A broken render can no longer reach the Library: it either passes the\ngate or the run fails. Upgrading the paper is now one deliberate command.</li>\n<li><strong>Next.</strong> Optionally regenerate figures in the same run once the <code>atlas-distill</code>\nJSON inputs are present in CI.</li>\n</ul>\n<h2 id=\"2026-05-18-fix-the-broken-working-paper-pdf\">2026-05-18 — Fix the broken working-paper PDF</h2><ul>\n<li><strong>Why.</strong> The linked <code>/immi_paper.pdf</code> was the stale <code>immi-paper-local.pdf</code> build: the\nabstract contained <strong>raw, unrendered LaTeX</strong> (<code>\\noindent\\textbf{Purpose:}</code>, <code>&lt;!--MATH0--&gt;</code>,\n<code>\\texttt{atlas-distill}</code>) and mojibake separators — an unreadable artifact.</li>\n<li><strong>What.</strong> Diagnosed via <code>pdftotext</code>: the served file (1.92 MB, Apr 29) leaked LaTeX\nmarkup and replacement characters. <code>paper/immi-paper-latest.pdf</code> (1.14 MB, May 5) is\nthe correct render — 0 LaTeX leaks, 0 mojibake, full references [1]–[29], and current\nscience (includes the d-band sample-size-confounder result). No LaTeX engine in this\nenvironment to recompile the 1-day-newer <code>.tex</code>, so swapped in the clean <code>-latest</code>\nbuild as <code>library-site/src/immi_paper.pdf</code>.</li>\n<li><strong>Results.</strong> The working paper now renders as a proper paper, ~785 KB smaller. Verified the\nbuilt <code>dist/immi_paper.pdf</code> has 0 raw-LaTeX leaks.</li>\n<li><strong>Next.</strong> Rebuild from <code>paper/immi-paper.tex</code> via <code>make</code> in a LaTeX environment if the\none-day-newer source has changes worth shipping; wire the paper build into CI so a\nbroken local PDF can&#39;t be the served one again.</li>\n</ul>\n<h2 id=\"2026-05-18-remove-entity-graph-fix-callout-filter-alignment\">2026-05-18 — Remove Entity Graph; fix callout/filter alignment</h2><ul>\n<li><strong>Why.</strong> The Entity Graph (force-graph) was unwanted weight, and the status-filter\npills and the paper/featured callouts hugged the left edge while the rest of the\npage is a centered 720px column — they used hardcoded inline <code>margin:0 16px</code> that\noverrode the column&#39;s <code>margin:0 auto</code>.</li>\n<li><strong>What.</strong> Deleted the Entity Graph end to end: the topbar button and <code>&lt;dialog&gt;</code> from\n<code>index.html</code>, the <code>force-graph</code> <code>&lt;script&gt;</code>, the whole graph section in <code>app.js</code>\n(~150 lines incl. the CDN fallback), the SW precache entry, the <code>build.js</code>\nvendoring step, the graph CSS, and the graph i18n strings. Added real <code>.status-filter</code>\n/ <code>.callout</code> / <code>.callout-box</code> classes that use the same <code>max-width:720px; margin:0 auto; padding:0 20px</code> container as <code>.shelf</code>/<code>.hero</code>, and removed the conflicting\ninline margins; callout text now uses theme variables instead of hardcoded <code>#fff</code>.</li>\n<li><strong>Results.</strong> No graph code, no <code>/vendor/</code>, no force-graph in the build; <code>app.js</code>\nsyntax-clean. The filter pills and both callouts now align flush with the shelves in\nevery theme.</li>\n<li><strong>Next.</strong> Optionally drop the unused <code>force-graph</code> dependency from <code>package.json</code>\n(left in place now to keep <code>npm ci</code> lockfile-in-sync; needs a lockfile regen).</li>\n</ul>\n<h2 id=\"2026-05-18-references-amp-lineage-shelf\">2026-05-18 — References &amp; Lineage shelf</h2><ul>\n<li><strong>Why.</strong> The corpus referenced ~35 external works (the IMMI <code>references.bib</code>) but a\nreader had nowhere to see the intellectual lineage — what we build on and why.</li>\n<li><strong>What.</strong> Authored <code>docs/references.md</code>: an annotated bibliography of all 35 works,\norganized by thread (sloppy-model theory → potentials → Simpson&#39;s/ecological fallacy →\nmeta-analysis → benchmark infra), each with citation, DOI, and a one-line <em>why we cite\nit</em>. Added a <strong>References &amp; Lineage</strong> shelf and moved the existing literature review\ninto it so external papers live in one place.</li>\n<li><strong>Results.</strong> 47 articles, 12 shelves. The program&#39;s lineage is now legible: each cited\nwork is tied to the load it bears (e.g. Mao et al. → why the ribbon should transfer to\nMLIPs; Pearl → the Lean Simpson&#39;s-paradox proof).</li>\n<li><strong>Next.</strong> Keep <code>references.md</code> in sync with <code>paper/references.bib</code> as the paper&#39;s\nbibliography grows.</li>\n</ul>\n<h2 id=\"2026-05-18-phase-2b-reader-side-status-filter\">2026-05-18 — Phase 2b: reader-side status filter</h2><ul>\n<li><strong>Why.</strong> Phase 2 shipped the status <em>badge</em> but you still could not <em>browse</em> by\nlifecycle stage — the named gap. A thinking surface should let you ask &quot;show me only\nwhat we refuted&quot; in one click.</li>\n<li><strong>What.</strong> Added a status filter bar on the home page: an <code>All · N</code> chip plus one\ncolor-coded chip per status actually present in the corpus, each with a live count.\nClicking filters every shelf to that status (spanning shelves, not just Conjectures —\ne.g. <code>proven</code> surfaces the Formal Proof Ledger), hides empty shelves and blurbs, and\n<code>All</code> restores. Pure client state, no manifest change.</li>\n<li><strong>Results.</strong> <code>supported → 3</code>, <code>refuted → 2</code>, <code>self-corrected → 1</code>, <code>proven → 1</code>,\n<code>open → 2</code>, <code>All → 46</code>. The refutations are now one click from the front door.</li>\n<li><strong>Next.</strong> Generate the per-hypothesis entries from the live closure records so the\nledger updates itself as research lands (the remaining Phase 2 follow-up).</li>\n</ul>\n<h2 id=\"2026-05-18-phase-2-the-corpus-becomes-a-ledger-tier-1-tier-2\">2026-05-18 — Phase 2: the corpus becomes a ledger (Tier 1 + Tier 2)</h2><ul>\n<li><strong>Why.</strong> The Library had the narrative (changelog) and the reports, but the actual\nscientific ledger — hypotheses, their status, the proofs and confounders behind them —\nexisted only as prose. A thinking surface needs the science browsable by <em>where it\nstands</em>, not just by topic.</li>\n<li><strong>What.</strong> Added <code>status</code> and <code>group</code> as first-class catalog/manifest/reader axes with a\ncolored lifecycle badge (proposed / supported / open / refuted / self-corrected /\nproven). Authored <strong>Tier 1</strong>: a Conjectures &amp; Proofs shelf (the hypothesis ledger +\n8 per-hypothesis entries, each Claim/Evidence/Confounder/Formal-cross-check/Next), a\nPartnerships shelf (MIIT-67 mapping under the public/gated convention), and a Formal\nProof Ledger mapping claims to Lean verdicts. Authored <strong>Tier 2</strong>: Data &amp; Provenance,\nMethodology (matched-n / contamination-gating / ecological-fallacy), and Reproduce Our\nResults.</li>\n<li><strong>Results.</strong> Library is now 46 articles across 11 shelves. The refutations\n(d-band, MEAM-2D) and the BCC/FCC self-correction are as visible as the confirmations,\neach with the confounder named — the self-correction discipline is finally legible as\nstructure, not buried in narrative.</li>\n<li><strong>Next.</strong> Reader-side status <em>filtering</em> (the badge ships now; faceted filter is the\nnext increment); generate the per-hypothesis entries from the live closure records so\nthe ledger updates as research lands.</li>\n</ul>\n<h2 id=\"2026-05-18-phase-1b-the-deploy-was-green-but-the-site-never-changed\">2026-05-18 — Phase 1b: the deploy was green but the site never changed</h2><ul>\n<li><strong>Why.</strong> After Phase 1 merged, the Cloud Build went green yet <code>library.lupine.science</code>\nstill showed no content and a white-square graph. &quot;Green build, stale site&quot; is the most\ndeceptive deploy failure there is.</li>\n<li><strong>What.</strong> Traced it past the domain and the direct <code>run.app</code> URL — both served the\nApril-28 build. <code>gcloud run services describe</code> showed the smoking gun: the <code>library-site</code>\nservice had <strong>traffic pinned 100% to revision <code>library-site-00013-kfj</code></strong>. Every deploy\nsince (we were at <code>00027</code>) created a healthy new revision that received <strong>0% traffic</strong>.\nVerified <code>00027</code> served the correct build via a temporary <code>verify</code> traffic tag, then\nmigrated traffic with <code>--to-latest</code> (which also sets <code>latestRevision: true</code>, so future\ndeploys auto-route). Hardened <code>cloudbuild.yaml</code> with an explicit step 6\n<code>gcloud run services update-traffic --to-latest</code> so a re-pin can never silently hide a\ndeploy again.</li>\n<li><strong>Results.</strong> <code>library.lupine.science</code> now serves the new build live: 26 articles, 8\nshelves incl. Changelog &amp; Progress, <code>changelog.json</code> 200, force-graph 200, SW\n<code>KILL=k1</code>. The graph code/CSS in the recovered source was never broken — the white\nsquare was the same stale revision. Service traffic config is now\n<code>{latestRevision: true, percent: 100}</code>.</li>\n<li><strong>Next.</strong> Returning visitors still running the <em>old</em> cache-first service worker need one\nhard reload for the new SW to take over (the old SW predates the <code>KILL</code> token, so it\ncan&#39;t self-evict — only the new SW can). After that, the network-first + <code>KILL</code> design\nprevents recurrence. Then Phase 2.</li>\n</ul>\n<h2 id=\"2026-05-18-phase-1-unbreak-the-library-deploy-path-self-healing-sw\">2026-05-18 — Phase 1: unbreak the Library deploy path (self-healing SW)</h2><ul>\n<li><strong>Why.</strong> <code>library.lupine.science</code> showed &quot;no content ever since the Chinese\nexperiment.&quot; Diagnosis: the server is <em>not</em> the problem — it returns 24 articles and\n200s. The break is a <strong>frozen Cloud Run image</strong> plus a <strong>cache-first service worker</strong>\nthat pins returning visitors to a stale/empty manifest forever, with no in-repo deploy\npath to ship a fix (the workflow was deleted with the source).</li>\n<li><strong>What.</strong> Confirmed the recovered <code>app.js</code>/<code>i18n.js</code> already fall back to EN safely\n(no logic bug to fix). Switched the service worker&#39;s <code>/data/*.json</code> from cache-first to\n<strong>network-first</strong> (fresh content wins; cache is the offline fallback) and added a <code>KILL</code>\ntoken to the cache namespace so any future bad build self-evicts on activate. Recreated\nthe deploy as committed IaC: <code>.github/workflows/deploy-library-site.yml</code> driving the\nrecovered <code>cloudbuild.yaml</code>, which replaces the frozen image on the existing\n<code>library-site</code> Cloud Run service.</li>\n<li><strong>Results.</strong> Clean local build: 26 articles, 8 shelves, force-graph self-hosted,\nversion+KILL-stamped SW, served smoke-test all 200s. The deploy path is now a\npush-to-main workflow, not a hand-run gcloud command. Not yet deployed to the live\ndomain — that triggers on merge to <code>main</code>.</li>\n<li><strong>Next.</strong> Merge to <code>main</code> to restore <code>library.lupine.science</code>; verify the SW takes over\nfor previously-bricked visitors; then Phase 2 (status/group facets + corpus-generated\nConjectures &amp; Proofs / Partnerships shelves).</li>\n</ul>\n<h2 id=\"2026-05-18-revive-the-library-as-the-public-thinking-surface\">2026-05-18 — Revive the Library as the public thinking surface</h2><ul>\n<li><strong>Why.</strong> <code>library.lupine.science</code> had been dark since a half-finished bilingual (EN/中文)\nexperiment; its source was deleted from the repo and the Cloud Run service was frozen on a\nstale, contentless build. The research corpus (reports, hypotheses, proofs, partnerships)\nhad no single public home and no changelog, so recent work — e.g. wiring Phoenix — was not\npointable-to.</li>\n<li><strong>What.</strong> Recovered the full <code>library-site/</code> static-site generator from git\n(<code>54e61f3^</code>). Rewrote the root README for external readers. Created this changelog in a\nwhy/what/results/next format. Deferred multi-language (keeping the <code>{en, zh}</code>-shaped\ncatalog so the door stays open) to focus on <strong>organization</strong>: category, group, and status\nas first-class axes so the Library is a place to <em>think about</em> the work, not just read it.</li>\n<li><strong>Results.</strong> Source restored and inspected; the catalog→Markdown→reader pipeline is intact\nand the model is well-suited to a status-aware research ledger. No redeploy yet (Phase 0 is\nintentionally local-only and reversible).</li>\n<li><strong>Next.</strong> (1) Add <code>status</code> and <code>group</code> to catalog entries + build manifest + reader facets\nso hypotheses can be browsed by lifecycle stage. (2) Promote the hypothesis lifecycle and\npartnerships to public shelves generated from the corpus. (3) Fix i18n to fall back to EN\nsafely, then redeploy over the frozen service.</li>\n</ul>\n<h2 id=\"2026-05-16-wire-phoenix-evals-end-to-end\">2026-05-16 — Wire Phoenix evals end to end</h2><ul>\n<li><strong>Why.</strong> The research loop in <code>glim-think</code> produced no observability: the hourly evaluation\nworkflow and the Worker trace export were two disconnected halves, neither configured by\ndeploy automation, so 0/300 spans reached Phoenix Cloud. We could not measure whether the\nloop was actually improving hypotheses.</li>\n<li><strong>What.</strong> Identified the two-halves split (GitHub repo secrets for the eval workflow vs.\nwrangler secrets for the Worker exporter) and set both. Consolidated <code>glim-think</code> to a\nsingle AI-SDK-native LLM path, deleting a dead second path that produced no spans.</li>\n<li><strong>Results.</strong> Root cause confirmed: the Worker had no Phoenix secrets and silently used a\nno-op localhost exporter. Separately <em>proved</em> a hard infrastructure limit — a Cloudflare\nWorker cannot export OTLP directly to Phoenix Cloud; the CF edge black-holes the\nsubrequest with a fake <code>200</code>. The Phoenix key was valid all along.</li>\n<li><strong>Next.</strong> Stand up a GCP egress relay (mirroring <code>deploy-otlp-relay.yml</code>) so Worker spans\nreach Phoenix; then use the lifecycle-trace + scientific-throughput evals as the loop&#39;s\nfitness function.</li>\n</ul>\n<h2 id=\"2026-05-16-de-myopize-the-corpus-the-hyper-ribbon-is-not-an-artifact\">2026-05-16 — De-myopize the corpus (the hyper-ribbon is not an artifact)</h2><ul>\n<li><strong>Why.</strong> The error corpus was ~99.5% elastic-constant (C_ij) records. If the hyper-ribbon\nonly appeared in C_ij, it could be an artifact of one property family rather than a real\nfeature of potential error.</li>\n<li><strong>What.</strong> Recovered real lattice constants (a₀) from MLIP provenance for 45 records and\nforced a joint C_ij + a₀ manifold per fleet run.</li>\n<li><strong>Results.</strong> The hyper-ribbon <strong>survives</strong> on the joint manifold (participation ratio\n1.05–2.05) — it is not an elastic-constant artifact. E_coh and B₀ predictions still require\nthe external compute pipeline before they can join the manifold.</li>\n<li><strong>Next.</strong> Extend the compute pipeline (Phase-D recipes) to produce E_coh / B₀ so the\nmanifold spans four property families, then re-test ribbon stability.</li>\n</ul>\n<h2 id=\"2026-05-16-self-correction-the-bcc-fcc-quot-causal-shield-quot-was-contamination\">2026-05-16 — Self-correction: the BCC/FCC &quot;causal shield&quot; was contamination</h2><ul>\n<li><strong>Why.</strong> A dramatic result (BCC vs. FCC error correlation 0.90 vs. 0.04, a &quot;causal shield&quot;)\nwas too strong; strong results in a noisy corpus deserve suspicion before celebration.</li>\n<li><strong>What.</strong> Audited the records behind the effect and added an ingest guard plus an\nidempotent purge to fleet step 0.</li>\n<li><strong>Results.</strong> The effect was a ~1.5% data-contamination artifact (19 corrupt records).\nCorpus purged to 1231 records, gated at <code>|pred| &gt; 1500</code> / <code>≤ 0</code>. Honest residual: a modest\nBCC &gt; FCC tendency, no Cauchy relation. The real contribution is the B → C2 → C3′ → C4\nself-correction arc — the same matched-n method that refuted the d-band hypothesis.</li>\n<li><strong>Next.</strong> Treat self-correction as a publishable primitive: every refuted claim gets a\nchangelog entry and a hypothesis-shelf status of <code>refuted (by us)</code>, with the confounder\nnamed.</li>\n</ul>\n<hr>\n<h1 id=\"backfill-work-since-the-site-went-stale-2026-04-27-2026-05-16\">Backfill — work since the site went stale (2026-04-27 → 2026-05-16)</h1><p>The Library froze on the 2026-04-28 build. ~355 commits landed before it was revived.\nThese are the arcs that matter, reconstructed so the corpus reflects current reality.\nNot every commit — the ones that changed what we believe or what the system can do.</p>\n<h2 id=\"2026-05-17-phase-d-close-the-loop-with-real-physics\">2026-05-17 — Phase D: close the loop with real physics</h2><ul>\n<li><strong>Why.</strong> Every result above is computed from <em>predicted</em> properties already in the corpus.\nTo validate recipes and extend the manifold beyond C_ij/a₀ (to E_coh, B₀) we need to run\nreal LAMMPS, not trust the cache.</li>\n<li><strong>What.</strong> Shipped the Phase-D compute resolution lane: a WAF-resilient HTTP client\n(retry + backoff + jitter, browser UA), a committed NIST harness, and a resilient compute\ndeploy. Then ran real LAMMPS through it.</li>\n<li><strong>Results.</strong> Running real physics immediately surfaced <strong>3 real recipe/integration bugs</strong>\n(P0/P1) that the cached pipeline had masked — exactly the point of the lane. Compute path\nnow deploys and survives datacenter WAF blocks.</li>\n<li><strong>Next.</strong> Produce E_coh / B₀ from Phase-D recipes and fold them into the joint manifold;\nre-test hyper-ribbon stability across four property families.</li>\n</ul>\n<h2 id=\"2026-05-17-the-self-improving-eval-loop-evolver-spine\">2026-05-17 — The self-improving eval loop (Evolver spine)</h2><ul>\n<li><strong>Why.</strong> Phoenix gave us observability; the next step is <em>actuation</em> — a loop that reads\nits own eval results and improves the thing being measured. Without it, evals are a\ndashboard, not a kernel.</li>\n<li><strong>What.</strong> Built the eval-loop units end to end: Phoenix dataset-read + Experiments REST\nclient, an A/B oracle, the Evolver self-improving spine, the registry/provenance/\nregression-gate trio, and <code>/ops/experiment-generate</code> (the Evolver activation prereq).\nClosed the eval→routing loop so model selection is eval-aware.</li>\n<li><strong>Results.</strong> The loop&#39;s spine exists and is wired: hypothesis lifecycle traces are the\nsubstrate, Phoenix evals are the fitness function, the Evolver is the actuator. Autonomous\nactuation is deliberately narrow (prompts/rubrics/criteria); structural change stays\nPR-gated.</li>\n<li><strong>Next.</strong> Arm the Evolver on a live hypothesis; keep structural edits human-gated. This is\nthe long-term organizing principle — the hypothesis lifecycle, not the prompt, is the unit\nof optimization.</li>\n</ul>\n<h2 id=\"2026-05-04-05-hypothesis-closures-the-corpus-refutes-itself-correctly\">2026-05-04…05 — Hypothesis closures: the corpus refutes itself, correctly</h2><ul>\n<li><strong>Why.</strong> The hyper-ribbon and the BCC/FCC dichotomy were found in <em>classical</em> potentials.\nDo they transfer to foundation MLIPs? And do our exciting sub-findings survive scrutiny?</li>\n<li><strong>What.</strong> Ran the research-round loop: ingested MACE-MP-0, then CHGNet, then Orb-v3 on the\nIMMI elements; ran matched-n bootstrap tests on the d-band and MEAM anomalies.</li>\n<li><strong>Results.</strong><ul>\n<li><strong>Hyper-ribbon transfers:</strong> 14/15 IMMI elements stay on the hyper-ribbon when each\nfoundation MLIP is added. This is the genuinely surprising result — we did <em>not</em> expect\nclassical→MLIP transfer.</li>\n<li><strong>Au escapes</strong> the ribbon across MACE and CHGNet (confirmed); <strong>Ag escape refuted</strong>\n(CHGNet pulls it back); <strong>Fe</strong> is a persistent outlier invariant to LAM addition.</li>\n<li><strong>D-band hypothesis REFUTED</strong> (full-sample ρ=−0.02); the apparent signal was a\nsample-size confounder (ρ=−0.50 to −0.66), recovering only on the n≥3 subset (ρ=+0.52).</li>\n<li><strong>MEAM &quot;intrinsically 2D&quot; anomaly REFUTED</strong> by matched-n bootstrap (MEAM n=7 median\nPR=1.36 overlaps Tersoff PR=1.01) — same confounder flavor as d-band.</li>\n</ul>\n</li>\n<li><strong>Next.</strong> These become <code>refuted (by us)</code> entries on the Conjectures &amp; Proofs shelf\n(Phase 2), each with the confounder named. The self-correction <em>method</em> is the\ncontribution.</li>\n</ul>\n<h2 id=\"2026-05-15-16-one-llm-path-eval-aware-routing-ai-gateway\">2026-05-15…16 — One LLM path, eval-aware routing, AI Gateway</h2><ul>\n<li><strong>Why.</strong> <code>glim-think</code> had two LLM paths; the gateway one was dead (0/300 Phoenix spans),\nso half the telemetry was fiction and routing decisions were blind.</li>\n<li><strong>What.</strong> Consolidated to one AI-SDK-native path (eval-aware deep tier), deleted the dead\ngateway path, routed Workers AI through Cloudflare AI Gateway (hybrid — Zhipu/MiniMax stay\ndirect, Gateway rejects them), and made the per-model scorecard read live-path attribution\n(<code>ai.telemetry.functionId</code>).</li>\n<li><strong>Results.</strong> Telemetry is now truthful; the eval→routing loop selects models on real\nmeasured performance instead of a path that never executed.</li>\n<li><strong>Next.</strong> Let the scorecard drive the Evolver&#39;s model-selection actuation.</li>\n</ul>\n<h2 id=\"2026-05-06-18-atlas-view-streaming-render-polish-curated-gallery\">2026-05-06…18 — atlas-view: streaming, render polish, curated gallery</h2><ul>\n<li><strong>Why.</strong> The WebGPU explorer choked on large scenes and shipped 185 uncurated gallery\nentries; the manifold is only persuasive if it renders fast and looks right.</li>\n<li><strong>What.</strong> Progressive chunked GPU upload + within-frame streaming parse + <code>.glimbin</code>\nstreaming pipeline + cluster-splat LOD + device-tier atom caps; bond/shader polish (flat\n2-tone bonds, isotropic atom shader, killed light flicker/shimmer); rebuilt the gallery\nto a curated 18-entry set; added a pre-merge CI and a Playwright UI harness.</li>\n<li><strong>Results.</strong> Huge scenes stream instead of stalling; the gallery is curated and the test\nsuite is green (14/14) with reproducible NIST catalog + streaming smoke tests in CI.</li>\n<li><strong>Next.</strong> Visual-regression diffing (screenshots are captured but not yet diffed).</li>\n</ul>\n<hr>\n<h2 id=\"how-to-add-an-entry\">How to add an entry</h2><p>Append at the top of the newest section. Keep Why/What/Results/Next. Prefer naming the\nconfounder, the null result, or the limit you hit — those are the entries that compound.\nThis file is wired into the Library catalog under the <strong>Changelog &amp; Progress</strong> shelf.</p>\n"}