{"id":"key-findings","title":"Phonon Benchmarking — Key Findings","subtitle":"Executive summary of the most consequential phonon findings for Lupine.","category":"validation","tags":["summary","phonon"],"source":"articles/docs/KEY_FINDINGS_SUMMARY.md","lang":"en","words":949,"readMinutes":4,"toc":[{"depth":2,"text":"Executive Summary","id":"executive-summary"},{"depth":2,"text":"Critical Findings for GLIM Implementation","id":"critical-findings-for-glim-implementation"},{"depth":3,"text":"1. Phonon Accuracy as the Gold Standard for Potential Validation","id":"1-phonon-accuracy-as-the-gold-standard-for-potential-validation"},{"depth":3,"text":"2. Accuracy Hierarchy Across Potential Families","id":"2-accuracy-hierarchy-across-potential-families"},{"depth":3,"text":"3. Composition-Dependent Performance Variations","id":"3-composition-dependent-performance-variations"},{"depth":3,"text":"4. Displacement-Dependent Phonon Errors","id":"4-displacement-dependent-phonon-errors"},{"depth":3,"text":"5. Reference Data and Functional Sensitivity","id":"5-reference-data-and-functional-sensitivity"},{"depth":3,"text":"6. Fine-Tuning Delivers Substantial Improvements","id":"6-fine-tuning-delivers-substantial-improvements"},{"depth":3,"text":"7. Dynamical Stability Prediction Challenges","id":"7-dynamical-stability-prediction-challenges"},{"depth":3,"text":"8. Computational Resource Planning","id":"8-computational-resource-planning"},{"depth":3,"text":"9. Hierarchical Metrics Enable Multi-Level Diagnosis","id":"9-hierarchical-metrics-enable-multi-level-diagnosis"},{"depth":3,"text":"10. Community-Driven Benchmark Evolution","id":"10-community-driven-benchmark-evolution"},{"depth":2,"text":"GLIM Strategic Recommendations","id":"glim-strategic-recommendations"},{"depth":2,"text":"Expected Timeline and Deliverables","id":"expected-timeline-and-deliverables"},{"depth":2,"text":"File Location","id":"file-location"}],"html":"<blockquote>\n<p>⚠️ <strong>Stale / superseded summary.</strong> This is an extraction-process summary of the\nphonon report with dead <code>/sessions/...</code> paths. For the current, complete review, see\n<a href=\"#/read/phonon-benchmarking\"><code>docs/phonon_benchmarking_report.md</code></a>.</p>\n</blockquote>\n<h1 id=\"phonon-benchmarking-report-key-findings-relevant-to-glim\">Phonon Benchmarking Report: Key Findings Relevant to GLIM</h1><h2 id=\"executive-summary\">Executive Summary</h2><p>The Phonon Frequency Spectrum Benchmarking Deep Research report provides a comprehensive technical framework for assessing 23 interatomic potentials across 12,000 materials (276,000 phonon calculations total). The report establishes systematic methodologies, performance metrics, and computational strategies essential for the GLIM project.</p>\n<h2 id=\"critical-findings-for-glim-implementation\">Critical Findings for GLIM Implementation</h2><h3 id=\"1-phonon-accuracy-as-the-gold-standard-for-potential-validation\">1. Phonon Accuracy as the Gold Standard for Potential Validation</h3><p><strong>Key Insight:</strong> Phonon frequencies probe second-order energy derivatives, making them exponentially more sensitive to potential errors than energies/forces (which depend on zeroth/first derivatives).</p>\n<ul>\n<li>Phonon benchmarks reveal &quot;force-constant collapse&quot; in neural network potentials: models can predict reasonable forces yet severely inaccurate force constants</li>\n<li>PhononBench evaluation of 108,843 AI-generated structures: only 25.83% dynamically stable</li>\n<li>Phonon validation is essential before deploying potentials for thermodynamic property prediction and materials discovery</li>\n</ul>\n<p><strong>GLIM Application:</strong> Phonon errors directly translate to property errors (~1% frequency error → ~1% entropy/free energy error, ~2% heat capacity error at room temperature). Explicit phonon testing prevents catastrophic failures in derived property predictions.</p>\n<h3 id=\"2-accuracy-hierarchy-across-potential-families\">2. Accuracy Hierarchy Across Potential Families</h3><p><strong>Clear Performance Ranking:</strong></p>\n<ul>\n<li>Universal MLIPs: 2–10 meV typical accuracy</li>\n<li>Fine-tuned MLIPs: &lt;2 meV (sub-meV achievable)</li>\n<li>Classical potentials: 20–100+ meV for general systems</li>\n</ul>\n<p><strong>Top Performers:</strong></p>\n<ul>\n<li><strong>ORB v3:</strong> Highest ranking in uMLIP benchmark, ~10× more efficient than OMat24</li>\n<li><strong>SevenNet-0:</strong> Top-tier accuracy with favorable computational scaling</li>\n<li><strong>MatterSim:</strong> DFT-level accuracy across &gt;10,000 materials</li>\n<li><strong>OMat24:</strong> Top-tier accuracy but drastically steeper scaling (CPU cost)</li>\n<li><strong>eqV2-M (fine-tuned):</strong> 0.174 log(W/m·K) MAE for thermal conductivity</li>\n</ul>\n<p><strong>GLIM Implication:</strong> Pareto frontier analysis enables optimal potential selection based on accuracy vs. computational cost constraints.</p>\n<h3 id=\"3-composition-dependent-performance-variations\">3. Composition-Dependent Performance Variations</h3><p><strong>Systematic Accuracy Pattern:</strong>\nMain-group compounds &gt; Transition metals &gt; Heavy elements/Rare earths</p>\n<p><strong>Challenging Systems:</strong></p>\n<ul>\n<li>Van der Waals materials: Large errors for PBE-trained models without dispersion corrections</li>\n<li>Systems with H + heavy elements: Unusual coordination patterns difficult to model</li>\n<li>Late transition metal oxides: Strong correlation effects not captured</li>\n</ul>\n<p><strong>GLIM Strategy:</strong> Stratified benchmarking by crystal system, chemistry (binary/ternary/quaternary), and bonding type enables identification of hidden failure modes and targeted improvements.</p>\n<h3 id=\"4-displacement-dependent-phonon-errors\">4. Displacement-Dependent Phonon Errors</h3><p><strong>Critical Discovery:</strong> ORB and OMat show dramatic MAE increase at small displacements (attributed to direct force output vs. Hessian-based approaches).</p>\n<p><strong>Implication:</strong> Displacement magnitude selection (standard 0.01–0.03 Å) is architecture-dependent. Quality control requires validation of force constant behavior across displacement ranges.</p>\n<h3 id=\"5-reference-data-and-functional-sensitivity\">5. Reference Data and Functional Sensitivity</h3><p><strong>JARVIS-DFT Database:</strong></p>\n<ul>\n<li>~90,000 materials with ~17,400 having complete phonon data</li>\n<li>Uses vdW-DF-optB88 functional (not PBE)</li>\n<li>Functional-induced frequency shifts: 2–6% (dense solids) to 20%+ (van der Waals)</li>\n</ul>\n<p><strong>GLIM Consideration:</strong> Functional mismatch between PBE-trained models and optB88 reference creates systematic bias. Recommendation: functional-corrected metrics or materials-specific comparison subsets.</p>\n<h3 id=\"6-fine-tuning-delivers-substantial-improvements\">6. Fine-Tuning Delivers Substantial Improvements</h3><p><strong>Phonon Force Constant Tuning (PFT):</strong></p>\n<ul>\n<li>55% average improvement in phonon properties</li>\n<li>Sub-meV errors achievable (ω_max MAE ~0.9 meV, S_vib MAE ~11 J/mol·K)</li>\n<li>Cost-effective: 60–140 GPU-hours training investment</li>\n<li>Transfer learning benefit: improves anharmonic properties (thermal conductivity κ: 0.446 → 0.306 log(W/m·K) SRME, 31% improvement)</li>\n</ul>\n<p><strong>GLIM Application:</strong> Fine-tuning pathways enable rapid accuracy optimization for top-performing potentials. The transfer from harmonic Hessian training to anharmonic properties suggests curvature supervision captures essential physics beyond harmonic regime.</p>\n<h3 id=\"7-dynamical-stability-prediction-challenges\">7. Dynamical Stability Prediction Challenges</h3><p><strong>Key Result:</strong> 25.83% baseline stability rate in AI-generated structures (high class imbalance).</p>\n<p><strong>Metric Interpretation:</strong></p>\n<ul>\n<li>Random guessing: 25% accuracy</li>\n<li>Perfect prediction: 100% accuracy</li>\n<li>Skill scores (improvement over baseline) enable fair comparison</li>\n<li>MatterSim achieves ~95% true-positive rate with validated false-positive/negative characterization</li>\n</ul>\n<p><strong>GLIM Use Case:</strong> Potentials must reliably identify imaginary modes for structure screening in materials discovery pipelines.</p>\n<h3 id=\"8-computational-resource-planning\">8. Computational Resource Planning</h3><p><strong>Efficiency Metrics:</strong></p>\n<ul>\n<li>Classical potentials: O(N) scaling, fastest per-calculation</li>\n<li>Efficient MLIPs (ORB, MACE): 10–100× speedup vs. DFT full workflow</li>\n<li>Large architectures (OMat24): Drastically steeper scaling than MACE/SevenNet</li>\n</ul>\n<p><strong>Cost Estimation for GLIM:</strong></p>\n<ul>\n<li>Base estimate: ~10⁶ GPU-hours for complete 23 × 12,000 benchmark</li>\n<li>With contingency (+50% for failures/restarts): ~1.5 × 10⁶ GPU-hours</li>\n<li>Parallelization strategy: GPU acceleration for NN inference, CPU clusters for classical potentials</li>\n</ul>\n<p><strong>Optimization Strategy:</strong></p>\n<ul>\n<li>Phase 1: All 23 potentials on 1,000-material representative subset (statistical power &amp; ranking)</li>\n<li>Phase 2: Top 10 potentials on full 12,000-material set (definitive comparison)</li>\n<li>Specialized analysis (fine-tuning, failure modes) on priority subsets</li>\n</ul>\n<h3 id=\"9-hierarchical-metrics-enable-multi-level-diagnosis\">9. Hierarchical Metrics Enable Multi-Level Diagnosis</h3><p><strong>Primary Metrics (Global Ranking):</strong></p>\n<ul>\n<li>Frequency MAE/RMSE across all q-points/branches</li>\n<li>Maximum frequency error (ω_max)</li>\n<li>Stability prediction accuracy (F1, ROC-AUC)</li>\n</ul>\n<p><strong>Secondary Metrics (Pattern Analysis):</strong></p>\n<ul>\n<li>PDOS Wasserstein distance and KL divergence</li>\n<li>Thermodynamic properties: S_vib, Helmholtz F, C_V accuracy</li>\n<li>Band-resolved MAE identifies mode-specific weaknesses</li>\n</ul>\n<p><strong>Diagnostic Metrics (Improvement Guidance):</strong></p>\n<ul>\n<li>Composition/structure-dependent error patterns</li>\n<li>Failure mode clustering</li>\n<li>Spectral feature accuracy (acoustic vs. optical, soft modes)</li>\n</ul>\n<h3 id=\"10-community-driven-benchmark-evolution\">10. Community-Driven Benchmark Evolution</h3><p><strong>Framework Design:</strong></p>\n<ul>\n<li>Annual major releases with new potentials, expanded materials, refined metrics</li>\n<li>Steering committee with academic, national lab, industry representation</li>\n<li>Specialized spin-off benchmarks for specific applications (battery materials, nuclear fuels, quantum materials)</li>\n<li>Open data release: 276,000 calculations with clear documentation</li>\n<li>Containerized code for portability and reproducibility</li>\n</ul>\n<p><strong>Ultimate Goal:</strong> Living benchmark that continuously improves with community contribution, enabling sustained progress in interatomic potential development.</p>\n<h2 id=\"glim-strategic-recommendations\">GLIM Strategic Recommendations</h2><ol>\n<li><p><strong>Adopt stratified sampling:</strong> By crystal system, chemistry type, bonding character to ensure balanced coverage and reveal hidden failure modes.</p>\n</li>\n<li><p><strong>Implement functional sensitivity analysis:</strong> Systematic PBE-PBEsol-optB88 comparisons on representative subsets to quantify functional bias.</p>\n</li>\n<li><p><strong>Prioritize fine-tuning:</strong> Top performers benefit from Hessian-targeted training with modest computational investment.</p>\n</li>\n<li><p><strong>Multi-level metric reporting:</strong> Provide frequency-level details (band structure), PDOS-level summaries, and property-level impacts simultaneously.</p>\n</li>\n<li><p><strong>Composition-specific skill scores:</strong> Report baseline-corrected metrics stratified by chemistry to enable fair multi-material comparison.</p>\n</li>\n<li><p><strong>Documentation of failure modes:</strong> Catalog systematic error patterns (displacement-dependent, mode-specific, chemistry-specific) to guide potential development.</p>\n</li>\n<li><p><strong>Benchmark versioning:</strong> Maintain backward compatibility while enabling evolving methodology through version tracking.</p>\n</li>\n</ol>\n<h2 id=\"expected-timeline-and-deliverables\">Expected Timeline and Deliverables</h2><ul>\n<li>Phase 1 (1000 materials, 23 potentials): Initial ranking &amp; statistical power assessment</li>\n<li>Phase 2 (12,000 materials, top 10 potentials): Comprehensive comparison</li>\n<li>Specialized analyses: Fine-tuning opportunities, failure mode deep dives</li>\n<li>Final deliverable: 276,000 phonon calculations + complete analysis framework</li>\n</ul>\n<h2 id=\"file-location\">File Location</h2><p>Full report: <code>/sessions/friendly-gracious-hamilton/breadth_exploration/dr_reports/phonon_benchmarking_report.md</code></p>\n"}