{"id":"research-evolution","title":"The Loop That Caught Itself","subtitle":"Research evolution report — how the corpus refuted its own exciting findings.","category":"changelog","tags":["self-correction","research-evolution","narrative"],"source":"articles/docs/research_evolution_2026_05_05.md","lang":"en","words":1837,"readMinutes":8,"toc":[{"depth":2,"text":"TL;DR","id":"tl-dr"},{"depth":2,"text":"§1 — The cycle, in one sentence","id":"1-the-cycle-in-one-sentence"},{"depth":2,"text":"§2 — The three stages, with concrete artifacts","id":"2-the-three-stages-with-concrete-artifacts"},{"depth":3,"text":"§2.1 Organize","id":"2-1-organize"},{"depth":3,"text":"§2.2 Harden","id":"2-2-harden"},{"depth":3,"text":"§2.3 Evaluate","id":"2-3-evaluate"},{"depth":2,"text":"§3 — Idea evolution, round by round","id":"3-idea-evolution-round-by-round"},{"depth":3,"text":"§3.1 Round 1 (2026-05-02) — hypcrossstylepc1universal","id":"3-1-round-1-2026-05-02-hypcrossstylepc1universal"},{"depth":3,"text":"§3.2 Round 1.5 (2026-05-02) — hyprankprscaling","id":"3-2-round-1-5-2026-05-02-hyprankprscaling"},{"depth":3,"text":"§3.3 Round 2 (2026-05-03) — hypglimmlipvalue","id":"3-3-round-2-2026-05-03-hypglimmlipvalue"},{"depth":3,"text":"§3.4 Round 3 (2026-05-03) — hyptop3lamdiagnostics","id":"3-4-round-3-2026-05-03-hyptop3lamdiagnostics"},{"depth":3,"text":"§3.5 Round 3-closure (2026-05-04) — hypalignmentdband","id":"3-5-round-3-closure-2026-05-04-hypalignmentdband"},{"depth":3,"text":"§3.6 Round 4 (2026-05-04) — h4mlipinvariance + hyptop3lamdiagnostics","id":"3-6-round-4-2026-05-04-h4mlipinvariance-hyptop3lamdiagnostics"},{"depth":3,"text":"§3.7 Round 5a (2026-05-05) — Au reconciliation","id":"3-7-round-5a-2026-05-05-au-reconciliation"},{"depth":3,"text":"§3.8 Round 5b (2026-05-05) — hypmlipalignmenttest","id":"3-8-round-5b-2026-05-05-hypmlipalignmenttest"},{"depth":3,"text":"§3.9 Round 5c (2026-05-05) — hypmeamanomaly","id":"3-9-round-5c-2026-05-05-hypmeamanomaly"},{"depth":2,"text":"§4 — The convergence demonstration","id":"4-the-convergence-demonstration"},{"depth":3,"text":"§4.1 d-band closure (2026-05-04)","id":"4-1-d-band-closure-2026-05-04"},{"depth":3,"text":"§4.2 MEAM bootstrap (2026-05-05)","id":"4-2-meam-bootstrap-2026-05-05"},{"depth":3,"text":"§4.3 The system property","id":"4-3-the-system-property"},{"depth":2,"text":"§5 — The X-scale, Y-iteration projection","id":"5-the-x-scale-y-iteration-projection"},{"depth":3,"text":"§5.1 Axis X — corpus scale","id":"5-1-axis-x-corpus-scale"},{"depth":3,"text":"§5.2 Axis Y — iteration scale","id":"5-2-axis-y-iteration-scale"},{"depth":3,"text":"§5.3 Self-correction rate","id":"5-3-self-correction-rate"},{"depth":2,"text":"§6 — What we will have achieved","id":"6-what-we-will-have-achieved"},{"depth":2,"text":"§7 — Companion documents","id":"7-companion-documents"}],"html":"<blockquote>\n<p>⚠️ <strong>Historical snapshot — some claims later re-audited.</strong> This report records the\n2026-05-05 state of the autonomous research loop. Empirical claims such as &quot;14/15\nelements stay on the hyper-ribbon&quot; were later Born-screened and placed under re-audit;\nsee <a href=\"#/read/conjecture-ledger\"><code>docs/conjectures/ledger.md</code></a> and\n<a href=\"#/read/changelog\"><code>CHANGELOG.md</code></a> for current status.</p>\n</blockquote>\n<h1 id=\"the-loop-that-caught-itself-lupine-research-evolution-report\">The Loop That Caught Itself — Lupine Research Evolution Report</h1><p><strong>Date:</strong> 2026-05-05\n<strong>Author:</strong> A. Welcing\n<strong>Companion route:</strong> <a href=\"https://lupine.science/evolution\"><code>/evolution</code></a> on the public site\n<strong>Companion ledger:</strong> <a href=\"https://glim-think-v1.aw-ab5.workers.dev/hypotheses\"><code>https://glim-think-v1.aw-ab5.workers.dev/hypotheses</code></a></p>\n<hr>\n<h2 id=\"tl-dr\">TL;DR</h2><p>Five rounds of autonomous research over four days produced a working demonstration of a three-stage cycle: <strong>organize information → harden logic → evaluate ideas</strong>. The headline finding is not any single result but the system property: <strong>the same matched-n bootstrap method caught two independent statistical artifacts in two days, in different scientific domains</strong>. That is the harden stage doing its job repeatedly, and it is what makes the loop self-correcting rather than a fancier version of cherry-picking.</p>\n<p>This document is the long-form record of how the ideas moved, why the matched-n bootstrap is now load-bearing, and what the X-scale, Y-iteration projection looks like under BigQuery + GCP.</p>\n<hr>\n<h2 id=\"1-the-cycle-in-one-sentence\">§1 — The cycle, in one sentence</h2><p>Every round of work on this project is a loop:</p>\n<ol>\n<li><strong>Organize information</strong> into a typed corpus keyed to a hypothesis_id.</li>\n<li><strong>Harden the logic</strong> against statistical artifacts before any claim is advanced.</li>\n<li><strong>Evaluate ideas</strong> by moving them through a strict hypothesis lifecycle with a gate.</li>\n</ol>\n<p>Each stage owns one job. The interesting behavior — the system catching its own mistakes — emerges from the <em>composition</em> of the three.</p>\n<hr>\n<h2 id=\"2-the-three-stages-with-concrete-artifacts\">§2 — The three stages, with concrete artifacts</h2><h3 id=\"2-1-organize\">§2.1 Organize</h3><p>Every record — a literature harvest, a LAMMPS benchmark, a foundation-MLIP elastic-constant sweep, a critique reply — lands as a typed row in either the local distill SQLite ledger or the public Cloudflare D1 mirror. As of 2026-05-05 the corpus spans 953 classical potentials, 18 functional-form families, three foundation MLIPs (MACE-MP-0, CHGNet, Orb-v3), 15 elements, 7,940 benchmark records, and 25 active hypotheses across proposed/testing/confirmed/refuted status.</p>\n<p>Key artifacts:</p>\n<ul>\n<li><code>POST /ingest/batch</code> — bulk record ingest. 45 records per MLIP × 3 MLIPs landed 2026-05-04.</li>\n<li><code>POST /admin/harvest</code> + <code>POST /admin/comprehend</code> — literature pipeline against arXiv + OpenAlex.</li>\n<li><code>POST /admin/manifold-recompute</code> — force re-PCA across all 15 elements after new ingest.</li>\n<li><code>lupine-distill::worker_sync</code> — best-effort auto-push of every local claim to the worker; HTTP failures log and continue, never block the local insert.</li>\n</ul>\n<h3 id=\"2-2-harden\">§2.2 Harden</h3><p>Every quantitative claim runs through deterministic statistical tests with no LLM in the loop: bootstrap CIs (10,000 iterations standard), permutation tests (5,000 shuffles), matched-n controls, and Spearman/Mann-Whitney pairings via the Causal Durable Object. The job at this stage is <strong>not to find effects — it is to kill artifacts</strong>.</p>\n<p>Key artifacts:</p>\n<ul>\n<li><code>POST /admin/d-band-analysis</code> — Causal-DO RPC. Spearman + permutation + bootstrap on cross-style PC1 alignment; pure deterministic, zero LLM cost.</li>\n<li><code>mlip_immi/meam_bootstrap.py</code> — 10,000 matched-n subsamples on MEAM error vectors against tersoff baseline; demonstrates the matched-n method in stand-alone Python.</li>\n<li><code>mlip_immi/cross_mlip_alignment.py</code> — pairwise cosines on unit error vectors across MACE/CHGNet/Orb.</li>\n<li><code>AutoHypothesisEvaluation</code> claim type — theorist auto-eval claims now fenced from PR-based hypotheses (see §3.5 Au reconciliation).</li>\n</ul>\n<h3 id=\"2-3-evaluate\">§2.3 Evaluate</h3><p>Hypotheses traverse a strict lifecycle: <code>proposed → testing → confirmed | refuted</code>, with status changes requiring <code>evidence_ids</code> attached and a Lean-readiness gate that refuses formalization until five boolean checks pass:</p>\n<ol>\n<li>confidence ≥ 0.85</li>\n<li>verdict stable across 3 rounds</li>\n<li>≥ 5 high-relevance insights</li>\n<li>no recent refutations</li>\n<li>narrative carries numerical anchors</li>\n</ol>\n<p>Synthesis claims close every round and queue the next set of hypotheses, each paired with a concrete experiment description on <code>/research/questions</code>. The <code>/admin/iterate</code> loop chases follow-up queries until the M2.7 reasoner emits no new ones — convergence detection saves real cost.</p>\n<p>Key artifacts:</p>\n<ul>\n<li><code>PATCH /hypotheses/{id}</code> — confidence + evidence_ids + status atomically updated.</li>\n<li><code>GET /admin/lean-status</code> — five-check formalization gate per hypothesis.</li>\n<li>Synthesis claims with <code>tested_hypotheses[]</code> + <code>verdicts{}</code> + <code>newly_proposed[]</code>.</li>\n<li><code>/admin/iterate</code> — M2.7 reason→harvest→comprehend→re-reason loop with convergence termination.</li>\n</ul>\n<hr>\n<h2 id=\"3-idea-evolution-round-by-round\">§3 — Idea evolution, round by round</h2><p>Nine canonical hypotheses, in order. Refutations are <strong>not dead ends</strong>: each reliably leaves behind a narrower, defensible claim.</p>\n<h3 id=\"3-1-round-1-2026-05-02-hyp-cross-style-pc1-universal\">§3.1 Round 1 (2026-05-02) — <code>hyp_cross_style_pc1_universal</code></h3><p>LLM counter-claim: PC1 is invariant to functional form across all elements. Test: pooled cross-style cosine 0.689 &lt; threshold 0.7. <strong>Refuted.</strong> Spinoff: element-level dichotomy emerged → <code>hyp_pc1_element_form_dichotomy</code> proposed at conf 0.85.</p>\n<h3 id=\"3-2-round-1-5-2026-05-02-hyp-rank-pr-scaling\">§3.2 Round 1.5 (2026-05-02) — <code>hyp_rank_pr_scaling</code></h3><p>Many-body rank → PR scaling. Test: Spearman ρ = -0.26, p = 0.42. <strong>Refuted.</strong> Spinoff: MEAM-as-outlier observed at full n → <code>hyp_meam_anomaly</code> proposed.</p>\n<h3 id=\"3-3-round-2-2026-05-03-hyp-glim-mlip-value\">§3.3 Round 2 (2026-05-03) — <code>hyp_glim_mlip_value</code></h3><p>Meta-claim: GLIM provides value for MLIP development. Test: 1/6 high-relevance insights. Auto-generated follow-up queries pulled ecological-fallacy literature from ecology, not MLIP-specific work. <strong>Gate blocked.</strong> Lesson promoted to feedback memory: concrete claims search; meta-claims do not.</p>\n<h3 id=\"3-4-round-3-2026-05-03-hyp-top3-lam-diagnostics\">§3.4 Round 3 (2026-05-03) — <code>hyp_top3_lam_diagnostics</code></h3><p>MACE-MP, CHGNet, Orb inherit hyper-ribbon; GLIM diagnostics catch MLIP failure modes. Pre-seeded with 4 manual harvests + 6 manual comprehends → 7/7 high-relevance, converged at round 2 of 3. Confidence trajectory: 0.45 → unchanged → 0.90 (post-empirical). <strong>Confirmed.</strong> Spinoff: empirical experiment unblocked — actually run the trio on the IMMI corpus.</p>\n<h3 id=\"3-5-round-3-closure-2026-05-04-hyp-alignment-d-band\">§3.5 Round 3-closure (2026-05-04) — <code>hyp_alignment_d_band</code></h3><p>Cross-style PC1 dichotomy explained by d-band fullness. Tests: ρ(d_count, alignment) = -0.02 on full sample (refuted); ρ(n_pairs, alignment) = -0.50 → -0.66 on subset (confounder confirmed). <strong>Refuted as stated, sample-size confounder found.</strong> Spinoff: <code>hyp_alignment_sample_size_artifact</code> confirmed at 0.90; matched-n method established as the harden-stage primary.</p>\n<h3 id=\"3-6-round-4-2026-05-04-h4-mlip-invariance-hyp-top3-lam-diagnostics\">§3.6 Round 4 (2026-05-04) — <code>h4_mlip_invariance</code> + <code>hyp_top3_lam_diagnostics</code></h3><p>Hyper-ribbon survives MLIP additions. Local inference of MACE-MP-0 → CHGNet → Orb-v3 on the IMMI 15-element corpus. Result: 14/15 elements still PR &lt; 2.0 across the trio (Fe lone outlier). <strong>Confirmed at conf 0.90 for both.</strong> Four new hypotheses queued: Au escape, Pt orthogonality, Fe persistent outlier, MLIP alignment test.</p>\n<h3 id=\"3-7-round-5a-2026-05-05-au-reconciliation\">§3.7 Round 5a (2026-05-05) — Au reconciliation</h3><p><code>hyp_au_specific_mlip_escape</code>. PR-based escape (1.02 → 1.41 with trio additions) vs rank-correlation auto-eval (within-style r ≥ 0.99, no Simpson attenuation). Apparent contradiction. <strong>Both stand:</strong> PR measures dimensional spread; rank-r measures monotonicity. Pearson r at n=3 with property magnitudes spanning ~5× is bounded near 1 by structural variance. <strong>Methodology hardened.</strong> The auto-eval framework is now fenced from PR-based hypotheses; n_records-per-style ≥ 9 threshold required before <code>supports_universal</code> verdicts can fire.</p>\n<h3 id=\"3-8-round-5b-2026-05-05-hyp-mlip-alignment-test\">§3.8 Round 5b (2026-05-05) — <code>hyp_mlip_alignment_test</code></h3><p>Element-form dichotomy extends to foundation MLIPs. Test: Spearman ρ = 0.19, p = 0.51 vs classical PC1 alignment (n = 15). <strong>Refuted (expected).</strong> Spinoff: noble-vs-refractory MLIP split discovered (group means 0.90 vs 0.22) → <code>hyp_noble_vs_refractory_mlip_split</code> (conf 0.75) and <code>hyp_pd_coherent_mlip_error_mode</code> (conf 0.70) queued.</p>\n<h3 id=\"3-9-round-5c-2026-05-05-hyp-meam-anomaly\">§3.9 Round 5c (2026-05-05) — <code>hyp_meam_anomaly</code></h3><p>MEAM PR=2.24 vs tersoff PR=1.01 at the same many-body rank — angular term proposed as mechanism. Test: 10,000 matched-n subsamples of MEAM at n=7. <strong>At matched n, MEAM median PR = 1.36, p05 = 1.04 — overlaps tersoff.</strong> <strong>Refuted as a comparison claim, sample-size confounder found.</strong> Same artifact pattern as the d-band closure. Spinoff: <code>hyp_meam_intrinsically_2d</code> (full-n bootstrap CI [1.58, 2.39] excludes 1-D null) preserves the narrower defensible claim at conf 0.80.</p>\n<hr>\n<h2 id=\"4-the-convergence-demonstration\">§4 — The convergence demonstration</h2><p>This is the load-bearing claim about the system. Two independent hypotheses, in different domains, both turned out to be <strong>sample-size confounders rather than physical phenomena</strong>. The same deterministic method caught both.</p>\n<h3 id=\"4-1-d-band-closure-2026-05-04\">§4.1 d-band closure (2026-05-04)</h3><div class=\"table-wrap\"><table><thead><tr>\n<th></th>\n<th></th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"\">Apparent finding</td>\n<td data-label=\"\">Closed-shell d10 → high cross-style alignment (ρ predicted ~+0.7 from physical theory)</td>\n</tr>\n<tr>\n<td data-label=\"\">Matched-n test</td>\n<td data-label=\"\">Restrict to n_pairs ≥ 3, controlling for sampling depth</td>\n</tr>\n<tr>\n<td data-label=\"\">What was revealed</td>\n<td data-label=\"\">On full sample, ρ(d_count, alignment) = -0.02 (null). The dichotomy is dominated by sample-size, not d-band</td>\n</tr>\n<tr>\n<td data-label=\"\">Residual surviving claim</td>\n<td data-label=\"\">On controlled subset, residual d-band signal recovers at ρ = +0.52, p = 0.087 — <code>hyp_dband_partial_signal</code></td>\n</tr>\n</tbody></table></div><h3 id=\"4-2-meam-bootstrap-2026-05-05\">§4.2 MEAM bootstrap (2026-05-05)</h3><div class=\"table-wrap\"><table><thead><tr>\n<th></th>\n<th></th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"\">Apparent finding</td>\n<td data-label=\"\">MEAM PR=2.24 vs tersoff PR=1.01 at the same many-body rank — angular term proposed as mechanism</td>\n</tr>\n<tr>\n<td data-label=\"\">Matched-n test</td>\n<td data-label=\"\">10,000 subsamples of MEAM at n=7 (matched to tersoff)</td>\n</tr>\n<tr>\n<td data-label=\"\">What was revealed</td>\n<td data-label=\"\">MEAM at n=7 median PR=1.36, p05=1.04. Tersoff observed sits at MEAM-at-n=7 5th percentile — barely distinguishable.</td>\n</tr>\n<tr>\n<td data-label=\"\">Residual surviving claim</td>\n<td data-label=\"\">MEAM full-n PR=2.07, bootstrap CI [1.58, 2.39] — narrower &quot;MEAM is sloppy in 2D at full n&quot; claim survives in <code>hyp_meam_intrinsically_2d</code></td>\n</tr>\n</tbody></table></div><h3 id=\"4-3-the-system-property\">§4.3 The system property</h3><p>The same operator (matched-n bootstrap, no LLM in the loop) refuted two hypotheses in different scientific domains using the same artifact pattern. That is not a one-off — it is the harden stage doing its job repeatedly. Every additional round adds another opportunity for the loop to catch itself.</p>\n<p>A measurable consequence: any IMMI-paper claim that compares PR or alignment across pair_style families must either restrict to comparable n or report a matched-n bootstrap. This is now methodology, not folklore.</p>\n<hr>\n<h2 id=\"5-the-x-scale-y-iteration-projection\">§5 — The X-scale, Y-iteration projection</h2><p>The point of demonstrating self-correction at small scale is to motivate the cost of running it at full scale. If a five-round, ~10⁴-record system already catches its own confounders, a ~10⁷-record, thousand-round system should produce a measurable self-correction rate — a KPI for autonomous science itself.</p>\n<h3 id=\"5-1-axis-x-corpus-scale\">§5.1 Axis X — corpus scale</h3><p><strong>Today (May 2026):</strong> 953 classical potentials × 18 families × 15 elements × 3 properties = 7,940 records. Three foundation MLIPs added 2026-05-04. ~10⁴ records.</p>\n<p><strong>Grand finale:</strong> Full Materials Project corpus (&gt;150k materials), all 600+ KIM/NIST potentials, the broader LAM landscape (M3GNet, SevenNet, GNoME, DPA-3, EquiformerV2, Allegro, NequIP), beyond elastics into phonons, defect energies, surface energies, vacancy formation, magnetic ground states. <strong>Order of magnitude: 10⁷–10⁸ records.</strong></p>\n<p><strong>Enabler:</strong> BigQuery as the structured ledger; Cloud Storage for artifacts; Pub/Sub for streaming ingest events from a fleet of MD runners.</p>\n<h3 id=\"5-2-axis-y-iteration-scale\">§5.2 Axis Y — iteration scale</h3><p><strong>Today:</strong> Five rounds across three days. Three closures landed in a single session (2026-05-05). Two methodological lessons (sample-size, n-threshold) promoted to feedback memory.</p>\n<p><strong>Grand finale:</strong> Thousands of rounds running concurrently via Cloud Run worker fleet. Each new MLIP architecture (or new data corpus, or new property class) becomes a round. Hypothesis half-life measured rather than visually inspected.</p>\n<p><strong>Enabler:</strong> Cloud Run + Cloud Tasks for the iterate worker pool; Vertex AI for the M2.7 reasoner backend; Looker dashboards on BigQuery for round throughput and convergence metrics.</p>\n<h3 id=\"5-3-self-correction-rate\">§5.3 Self-correction rate</h3><p><strong>Today:</strong> Two independent confounders caught with the same matched-n method, two days apart. Multiple LLM counter-claims refuted by the deterministic harden stage. The Lean-readiness gate has refused every formalization attempt so far.</p>\n<p><strong>Grand finale:</strong> Self-correction rate becomes a measurable property. Per 1,000 hypotheses, what fraction get refuted? Per 100 confounders, what fraction are caught before formalization? These become KPIs of the autonomous research system itself, reported on a public dashboard.</p>\n<p><strong>Enabler:</strong> Cloud Logging + BigQuery analytical views; the existing /admin/lean-status overview becomes a SQL query against persistent state.</p>\n<hr>\n<h2 id=\"6-what-we-will-have-achieved\">§6 — What we will have achieved</h2><p>If the loop runs at the projected scale without intervention beyond hypothesis seeding, the deliverable is the first scientific reasoning system whose <strong>self-correction rate is auditable in public</strong> and whose <strong>artifact-detection record is written down</strong> rather than buried in lab folklore.</p>\n<p>That is a different deliverable from &quot;an automated literature reviewer&quot; or &quot;a fancy prompt-engineering harness.&quot; It is a system whose meta-properties — refutation rate, confounder catch-rate, formalization-gate failure rate — are first-class observables. The IMMI manuscript is the first paper produced by it; the second will be the operating-system paper itself.</p>\n<hr>\n<h2 id=\"7-companion-documents\">§7 — Companion documents</h2><ul>\n<li><a href=\"https://lupine.science/evolution\"><code>/evolution</code></a> — public TanStack route mirroring this report.</li>\n<li><a href=\"https://lupine.science/process\"><code>/process</code></a> — the original operating report (rounds 1–3).</li>\n<li><a href=\"https://lupine.science/research\"><code>/research</code></a> — the working-paper summary.</li>\n<li><a href=\"./plans/grand_finale_gcp.md\" target=\"_blank\" rel=\"noopener\" class=\"ll-raw-source\"><code>docs/plans/grand_finale_gcp.md</code></a> — detailed BigQuery + GCP migration plan.</li>\n<li><a href=\"../paper/immi-paper.tex\"><code>paper/immi-paper.tex</code></a> — working-paper source (currently marked WORK IN PROGRESS).</li>\n<li>Public ledger: <code>https://glim-think-v1.aw-ab5.workers.dev/hypotheses</code></li>\n</ul>\n"}