{"id":"internal-science-program","title":"Internal Science Program","subtitle":"The research agenda: what we are trying to learn and how the loop pursues it.","category":"foundations","tags":["program","agenda","vision"],"source":"articles/docs/internal-science-program.md","lang":"en","words":782,"readMinutes":4,"toc":[{"depth":2,"text":"Current Scientific Center","id":"current-scientific-center"},{"depth":2,"text":"Fresh Local Run","id":"fresh-local-run"},{"depth":2,"text":"Claim Triage","id":"claim-triage"},{"depth":3,"text":"Promote","id":"promote"},{"depth":3,"text":"Quarantine","id":"quarantine"},{"depth":3,"text":"Retire Or Rewrite","id":"retire-or-rewrite"},{"depth":2,"text":"Next Internal Experiments","id":"next-internal-experiments"},{"depth":2,"text":"Why This Is Exciting","id":"why-this-is-exciting"}],"html":"<h1 id=\"internal-science-program\">Internal Science Program</h1><p>This is a private working map for the next evidence push. Do not treat it as a\npublic claim sheet. It is here to keep the research honest while the system\ngrinds.</p>\n<h2 id=\"current-scientific-center\">Current Scientific Center</h2><p>The strongest idea is not that one potential ranks above another. The strongest\nidea is that interatomic-potential errors can be studied as physical objects:\nthey have covariance spectra, symmetry structure, provenance, failure modes,\nand causal context.</p>\n<p>The current center of gravity is:</p>\n<ol>\n<li>Error vectors for elastic observables are strongly compressed.</li>\n<li>Benchmark heterogeneity is signal, not nuisance.</li>\n<li>Simpson-style claims need strict causal auditing before they are allowed back\ninto headline status.</li>\n<li>The next real battlefield is curvature: phonons, Hessians, force constants,\nand non-equilibrium configurations.</li>\n<li>Formalization should trail only stable claims with artifacts and falsification\nnotes.</li>\n</ol>\n<h2 id=\"fresh-local-run\">Fresh Local Run</h2><p>Command run privately into <code>.glim-runtime/science-runs/</code>:</p>\n<pre><code class=\"language-powershell\">cargo run --release --manifest-path ..\\..\\atlas-distill\\Cargo.toml --bin atlas-distill -- benchmark (Resolve-Path ..\\..\\nist_benchmark.csv) --full\n</code></pre>\n<p>Input: <code>nist_benchmark.csv</code></p>\n<p>Dataset summary:</p>\n<ul>\n<li>Entries: 386</li>\n<li>Materials: 10, namely Ag, Al, Au, Cr, Cu, Fe, Mo, Ni, V, W</li>\n<li>Potentials: 38</li>\n<li>Properties: C11, C12, C44, Ecoh, a0</li>\n<li>Completeness: 13.8%</li>\n</ul>\n<p>Manifold results from the private run:</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Potential</th>\n<th align=\"right\">Materials</th>\n<th align=\"right\">PR / 5</th>\n<th>PR CI</th>\n<th align=\"right\">First mode variance</th>\n<th>Strict ribbon</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Potential\">Ackland-1987</td>\n<td align=\"right\" data-label=\"Materials\">4</td>\n<td align=\"right\" data-label=\"PR / 5\">1.458</td>\n<td data-label=\"PR CI\">1.000-1.697</td>\n<td align=\"right\" data-label=\"First mode variance\">80.5%</td>\n<td data-label=\"Strict ribbon\">no</td>\n</tr>\n<tr>\n<td data-label=\"Potential\">Foiles-1986</td>\n<td align=\"right\" data-label=\"Materials\">4</td>\n<td align=\"right\" data-label=\"PR / 5\">1.685</td>\n<td data-label=\"PR CI\">1.000-1.972</td>\n<td align=\"right\" data-label=\"First mode variance\">74.6%</td>\n<td data-label=\"Strict ribbon\">no</td>\n</tr>\n<tr>\n<td data-label=\"Potential\">Zhou-2004</td>\n<td align=\"right\" data-label=\"Materials\">5</td>\n<td align=\"right\" data-label=\"PR / 5\">1.150</td>\n<td data-label=\"PR CI\">1.000-1.442</td>\n<td align=\"right\" data-label=\"First mode variance\">93.0%</td>\n<td data-label=\"Strict ribbon\">no</td>\n</tr>\n<tr>\n<td data-label=\"Potential\">Adams-1989</td>\n<td align=\"right\" data-label=\"Materials\">4</td>\n<td align=\"right\" data-label=\"PR / 5\">1.991</td>\n<td data-label=\"PR CI\">1.000-1.991</td>\n<td align=\"right\" data-label=\"First mode variance\">65.9%</td>\n<td data-label=\"Strict ribbon\">no</td>\n</tr>\n</tbody></table></div><p>Interpretation:</p>\n<ul>\n<li>Low-dimensional compression is real and strong in these slices.</li>\n<li>The stricter geometric-ribbon classifier currently rejects all four sparse 5D\nslices.</li>\n<li>That means the next claim should be &quot;compressed error subspace&quot; unless and\nuntil the geometric-sequence law is separately proven.</li>\n<li>The zero or near-zero trailing eigenvalues are partly a sparse-rank issue:\nfour or five materials cannot fully populate a five-observable covariance\nmatrix.</li>\n</ul>\n<p>Meta-analysis results from the same run:</p>\n<ul>\n<li>Fixed effects: pooled r = 0.9901, I2 = 96.1%</li>\n<li>Random effects: pooled r = 0.9945, I2 = 96.5%</li>\n<li>Random-effects prediction interval: 0.8251-0.9998</li>\n</ul>\n<p>Interpretation:</p>\n<ul>\n<li>Correlations are very high overall, but heterogeneity remains extreme.</li>\n<li>The right claim is not &quot;one pooled number is enough.&quot; It is &quot;the pooled number\nsurvives here, while the heterogeneity says group structure still matters.&quot;</li>\n</ul>\n<h2 id=\"claim-triage\">Claim Triage</h2><h3 id=\"promote\">Promote</h3><p>Compressed error subspace:</p>\n<ul>\n<li>Supported by FCC, BCC, and fresh NIST slices.</li>\n<li>Good next form: &quot;Participation ratios stay near 1-2 across elastic/statics\nobservable bundles, even when the ambient observable count is 5.&quot;</li>\n</ul>\n<p>Heterogeneity as diagnostic:</p>\n<ul>\n<li>Supported by high I2 across local meta-analysis runs.</li>\n<li>Good next form: &quot;Random-effects meta-analysis should be a standard\nbenchmark diagnostic, even when all subgroup correlations are positive.&quot;</li>\n</ul>\n<p>Curvature validation:</p>\n<ul>\n<li>Supported by the phonon report and by the conceptual gap between elastic\nconstants and full Hessian behavior.</li>\n<li>Good next form: &quot;A potential is not validation-complete until force-constant\nand phonon stability errors are measured.&quot;</li>\n</ul>\n<h3 id=\"quarantine\">Quarantine</h3><p>Simpson&#39;s paradox:</p>\n<ul>\n<li>Existing local artifacts disagree in emphasis.</li>\n<li><code>paradox_detection.json</code> says no Simpson sign reversal but flags ecological\nfallacy by reversal magnitude.</li>\n<li>Lean causal docs say empirical Simpson and ecological fallacy are both absent\nfor a separate embedded dataset.</li>\n<li>Next action: build a single paradox audit table keyed by dataset, grouping,\nx/y definition, pooled r, pooled-within r, sign reversal, and magnitude gap.</li>\n</ul>\n<p>Strict hyper-ribbon law:</p>\n<ul>\n<li>The fresh 386-entry run shows strong compression but strict classifier failure.</li>\n<li>Next action: split the theorem into two layers:<ul>\n<li>Low PR compression.</li>\n<li>Geometric eigenvalue law.</li>\n</ul>\n</li>\n</ul>\n<h3 id=\"retire-or-rewrite\">Retire Or Rewrite</h3><p>Any public wording that says &quot;Simpson&#39;s paradox proven&quot; should be treated as\noutdated until the unified causal audit says otherwise.</p>\n<p>Any public wording that says &quot;hyper-ribbon proven&quot; should specify which\ncriterion is meant. Low PR is not the same as a strict geometric spectrum.</p>\n<h2 id=\"next-internal-experiments\">Next Internal Experiments</h2><ol>\n<li><p>Causal audit matrix</p>\n<ul>\n<li>Run every paradox detector against every local dataset.</li>\n<li>Standardize x/y definitions.</li>\n<li>Output one table with no narrative.</li>\n</ul>\n</li>\n<li><p>Rank-aware manifold audit</p>\n<ul>\n<li>For each potential, record sample count, observable count, matrix rank,\nPR, PR CI, geometric residual CV, and strict-ribbon result.</li>\n<li>Flag any run where sample count &lt;= observable count.</li>\n</ul>\n</li>\n<li><p>Bulk/shear mode basis</p>\n<ul>\n<li>Transform C11, C12, C44 into bulk-like and shear-like coordinates.</li>\n<li>Test whether principal directions align with physical elastic modes.</li>\n<li>This is the real path toward spectral rigidity.</li>\n</ul>\n</li>\n<li><p>Phonon sentinel protocol</p>\n<ul>\n<li>Start with small displacement sweeps for Al, Cu, Ni, and Ag.</li>\n<li>Record Hessian/force-constant sensitivity before attempting broad MLIP\nbenchmarking.</li>\n<li>Gate with dynamic-stability classification, not just frequency MAE.</li>\n</ul>\n</li>\n<li><p>Lean gate preparation</p>\n<ul>\n<li>Formalize low-PR compression separately from geometric spectrum.</li>\n<li>Add an explicit &quot;insufficient rank&quot; theorem/guard for sparse covariance\nclaims.</li>\n<li>Keep Simpson as a refuted or quarantined claim until the audit resolves.</li>\n</ul>\n</li>\n</ol>\n<h2 id=\"why-this-is-exciting\">Why This Is Exciting</h2><p>The scientific path is becoming sharper. GLIM is not merely a validation runner;\nit is a machine for discovering the structure of model error. If that holds,\nthen the useful product is a living error atlas: which observables collapse,\nwhich modes stay stiff, which potential families fail by the same geometry, and\nwhich new experiments break the compression.</p>\n"}