{"id":"foundation-model-trust-layers","title":"Foundation Materials Models Need Trust Layers","subtitle":"A ranked next-round target memo: batteries, MPtrj, Ni defects, Fe magnetism, Au surfaces, and phonons as evidence-carrying canaries.","category":"changelog","tags":["mlip","materials-discovery","trust-layer","batteries","featured"],"source":"articles/docs/foundation-model-trust-layers.md","lang":"en","words":1028,"readMinutes":5,"toc":[{"depth":2,"text":"Human summary","id":"human-summary"},{"depth":2,"text":"The literature signal","id":"the-literature-signal"},{"depth":2,"text":"Property-specific trust contracts","id":"property-specific-trust-contracts"},{"depth":2,"text":"Ranked Lupine target lanes","id":"ranked-lupine-target-lanes"},{"depth":3,"text":"1. Li solid-ion and Li-channel dynamics","id":"1-li-solid-ion-and-li-channel-dynamics"},{"depth":3,"text":"2. MPtrj broad-DFT canary","id":"2-mptrj-broad-dft-canary"},{"depth":3,"text":"3. Ni vacancy and defect transport","id":"3-ni-vacancy-and-defect-transport"},{"depth":3,"text":"4. Fe magnetic MLIP failure audit","id":"4-fe-magnetic-mlip-failure-audit"},{"depth":3,"text":"5. Au and heavy noble-metal surfaces","id":"5-au-and-heavy-noble-metal-surfaces"},{"depth":3,"text":"6. Phonon / Hessian trust for semiconductors, vdW materials, and ionic oxides","id":"6-phonon-hessian-trust-for-semiconductors-vdw-materials-and-ionic-oxides"},{"depth":2,"text":"Recommended public artifact format","id":"recommended-public-artifact-format"},{"depth":2,"text":"Working-paper patches implied by this memo","id":"working-paper-patches-implied-by-this-memo"},{"depth":2,"text":"Bottom line","id":"bottom-line"}],"html":"<h1 id=\"foundation-materials-models-need-trust-layers-not-just-bigger-benchmarks\">Foundation Materials Models Need Trust Layers, Not Just Bigger Benchmarks</h1><p><strong>Status:</strong> proposed / publication-ready library memo\n<strong>Publication date:</strong> 2026-06-20\n<strong>Audience:</strong> materials labs, MLIP builders, and observers evaluating the Lupine evidence trail.</p>\n<h2 id=\"human-summary\">Human summary</h2><p>Universal machine-learning interatomic potentials and generative materials models have become good enough to change the substrate of materials discovery. The bottleneck is no longer only “can we generate or relax more candidates?” It is “which prediction is reliable enough to spend DFT, synthesis, or lab time on?”</p>\n<p>Lupine’s next public lane should frame itself as the correction and trust layer for that substrate. The system should not claim to have discovered a new material yet. It should publish evidence-carrying canaries: sealed candidate sets, paired baseline-vs-correction comparisons, property-specific guards, refusal logs, and kill criteria.</p>\n<h2 id=\"the-literature-signal\">The literature signal</h2><p>Recent public work points in the same direction:</p>\n<ul>\n<li>MatterSim and related universal atomistic models broaden the usable domain of MLIPs across elements, temperatures, and pressures.</li>\n<li>CHGNet, MACE, ORB, SevenNet, MatterSim, and other foundation potentials are now practical defaults for screening and relaxation workflows.</li>\n<li>GNoME and MatterGen-style systems make candidate generation abundant.</li>\n<li>Battery, phonon, pressure, zeolite, surface, and kinetics benchmarks show that good average force or energy error does not certify every property decision.</li>\n<li>Active-learning, uncertainty, evidential, and delta-correction papers increasingly treat “when to trust the model” as the central question.</li>\n</ul>\n<p>The Lupine article thesis should be simple: foundation models make screening cheap enough that correction, calibration, and evidence become the scarce layer.</p>\n<h2 id=\"property-specific-trust-contracts\">Property-specific trust contracts</h2><p>A single MLIP accuracy number is not a discovery certificate. Each material candidate should carry separate trust status for the property being used:</p>\n<div class=\"table-wrap\"><table><thead><tr>\n<th>Trust contract</th>\n<th>What it certifies</th>\n<th>What can break it</th>\n</tr>\n</thead><tbody><tr>\n<td data-label=\"Trust contract\">Relaxation trust</td>\n<td data-label=\"What it certifies\">relaxed geometry is physically plausible</td>\n<td data-label=\"What can break it\">hidden phase change, unstable tensor, force explosion</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Energy-ranking trust</td>\n<td data-label=\"What it certifies\">candidates are ordered correctly enough for screening</td>\n<td data-label=\"What can break it\">functional bias, stoichiometry bias, support mismatch</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Force/MD trust</td>\n<td data-label=\"What it certifies\">finite-temperature trajectory is usable</td>\n<td data-label=\"What can break it\">drift, rare-event failures, bad local curvature</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Phonon trust</td>\n<td data-label=\"What it certifies\">dynamical stability and second derivatives are credible</td>\n<td data-label=\"What can break it\">imaginary modes, displacement sensitivity, Hessian collapse</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Migration-barrier trust</td>\n<td data-label=\"What it certifies\">NEB / hop barriers are decision-relevant</td>\n<td data-label=\"What can break it\">wrong transition state, barrier ranking inversion</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Surface/catalysis trust</td>\n<td data-label=\"What it certifies\">adsorption and reaction surfaces are credible</td>\n<td data-label=\"What can break it\">bulk-training extrapolation, charge/coverage effects</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Pressure/temperature trust</td>\n<td data-label=\"What it certifies\">model survives thermodynamic extrapolation</td>\n<td data-label=\"What can break it\">out-of-domain pressure, phase-boundary instability</td>\n</tr>\n<tr>\n<td data-label=\"Trust contract\">Electronic-property trust</td>\n<td data-label=\"What it certifies\">structure is suitable for band/optical claims</td>\n<td data-label=\"What can break it\">MLIP geometry confidence does not imply electronic accuracy</td>\n</tr>\n</tbody></table></div><h2 id=\"ranked-lupine-target-lanes\">Ranked Lupine target lanes</h2><h3 id=\"1-li-solid-ion-and-li-channel-dynamics\">1. Li solid-ion and Li-channel dynamics</h3><p>This is the strongest discovery-facing wedge. The repo already contains a LiFePO4 local canary where a paired correction improves final position RMSE from about 0.131 Å to about 0.042 Å under a force guard. That is not a conductivity discovery, but it is exactly the kind of sealed paired result that can mature into a battery trust lane.</p>\n<p>Next gate: lock a public solid-ion-conductor reference subset such as Li3YCl6 or Li6PS5Cl, run MACE/CHGNet/ORB/SevenNet baselines and corrections, and score position, force, stress, energy drift, migration proxy, and refusal rate.</p>\n<p>Kill if the correction improves position while worsening force/stress, fails shared-checkpoint pairing, or cannot lock a redistributable reference.</p>\n<h3 id=\"2-mptrj-broad-dft-canary\">2. MPtrj broad-DFT canary</h3><p>The MPtrj canary is the best bridge beyond the controlled Ni fixture. The support-floor v2 cloud result reports eight paired comparisons, six improvements, two safe holds, and zero regressions. That is a useful evidence-contract result, not a universal model-improvement result.</p>\n<p>Next gate: replay the row-hybrid v3 policy in Cloud Run with force/stress/elastic rows promoted to first-class gates.</p>\n<p>Kill if the win remains energy-only while the article implies broader model improvement.</p>\n<h3 id=\"3-ni-vacancy-and-defect-transport\">3. Ni vacancy and defect transport</h3><p>Ni is a credibility anchor. The broad Ni paired-accuracy campaign rejected promotion: 25 measured pairs, zero improvements, ten regressions, and fifteen unchanged. That negative result is useful because the gate caught harm. A local Ni vacancy canary is promising but not yet externally locked.</p>\n<p>Next gate: lock vacancy formation, local relaxation, and migration references before promoting any defect-transport claim.</p>\n<p>Kill if bulk anchors regress or the local improvement fails cloud/shared-checkpoint replay.</p>\n<h3 id=\"4-fe-magnetic-mlip-failure-audit\">4. Fe magnetic MLIP failure audit</h3><p>Fe remains strategically important, but the old “persistent PR outlier” line is not currently citable after Born screening. The credible target is narrower: a magnetic mechanical-stability failure audit for Fe and related transition metals.</p>\n<p>Next gate: compare spin-aware and spin-agnostic references under Born screening, phonon stability, elastic stability, and coupling-aware nulls.</p>\n<p>Kill if matched screened inputs remove the effect or if the signal is explained by one invalid tensor.</p>\n<h3 id=\"5-au-and-heavy-noble-metal-surfaces\">5. Au and heavy noble-metal surfaces</h3><p>Au escape is a good hypothesis target, especially for catalyst-adjacent surfaces, but it remains open. The bulk pre-screening signal should become a surface/adsorbate canary with Ag/Pt/Pd controls.</p>\n<p>Next gate: low-index slabs and simple adsorbates with MACE/CHGNet/ORB/SevenNet, reporting scalar errors plus error-subspace geometry.</p>\n<p>Kill if ORB/SevenNet do not reproduce the effect or if a group-level noble-metal explanation beats Au-specificity.</p>\n<h3 id=\"6-phonon-hessian-trust-for-semiconductors-vdw-materials-and-ionic-oxides\">6. Phonon / Hessian trust for semiconductors, vdW materials, and ionic oxides</h3><p>Phonons are strategically important because they expose second-derivative failures that scalar energy/force metrics can hide. Current repo evidence is mostly protocol and review, so this should be framed as the next canary protocol unless pilot measurements are added.</p>\n<p>Next gate: a small sealed phonon subset across Si/Ge/diamond, MgO/Al2O3/SrTiO3, and graphite/MoS2/h-BN.</p>\n<p>Kill if reference functional mismatch dominates or displacement-size sensitivity changes the conclusion.</p>\n<h2 id=\"recommended-public-artifact-format\">Recommended public artifact format</h2><p>Every top-priority target should ship as an evidence-carrying candidate, not as a press release:</p>\n<ul>\n<li>material or class;</li>\n<li>model family;</li>\n<li>property target;</li>\n<li>baseline artifact;</li>\n<li>corrected artifact;</li>\n<li>reference source and license;</li>\n<li>support-domain signal;</li>\n<li>uncertainty/calibration signal;</li>\n<li>refusal decision;</li>\n<li>no-regression checks;</li>\n<li>kill criterion;</li>\n<li>LUPI or library route for inspection.</li>\n</ul>\n<h2 id=\"working-paper-patches-implied-by-this-memo\">Working-paper patches implied by this memo</h2><ol>\n<li>Replace any scalar “MLIP accuracy” language with property-specific trust contracts.</li>\n<li>Say explicitly that correction is decision-specific: barrier trust is not phonon trust, and geometry trust is not electronic-property trust.</li>\n<li>Treat generative discovery outputs as candidates that need evidence, not as discoveries by default.</li>\n<li>Promote batteries and catalysts only where the repo has sealed canaries or a concrete next gate.</li>\n<li>Keep construction/cement and broad catalyst discovery out of the headline until there is a source packet and a measured canary.</li>\n</ol>\n<h2 id=\"bottom-line\">Bottom line</h2><p>The next Lupine publication should not overclaim a new material. It should show the infrastructure that makes future material claims worth believing: a trust layer that catches negative transfer, refuses unsafe corrections, and publishes the evidence trail attached to each candidate.</p>\n"}