{"id":"lit-review-error-structure","title":"Error Structure in Interatomic Potentials — Literature Review","subtitle":"Sloppy models, Simpson’s paradox, and FCC error manifolds: a 30-reference synthesis.","category":"references","tags":["literature-review","sloppy","simpson","fcc"],"source":"articles/lit-review.md","lang":"en","words":2198,"readMinutes":10,"toc":[],"html":"<p>Literature review: error structure in interatomic</p>\n<p>potential predictions</p>\n<p>Classical force fields for metals exhibit structured, low-dimensional prediction errors that</p>\n<p>can be understood through the convergence of three distinct intellectual traditions: sloppy</p>\n<p>model theory from statistical physics, Simpson’s paradox from meta-analytic methodology,</p>\n<p>and systematic interatomic potential benchmarking. This review assembles the key</p>\n<p>references across these areas to support a paper showing that FCC elastic constant errors</p>\n<p>(C11, C12, C44) occupy a manifold with effective dimensionality ~1.66/3, while pooled BCC</p>\n<p>data produce a spurious correlation inversion attributable to Simpson’s paradox.</p>\n<p>Area 1: Sloppy model theory and the geometry of poorly</p>\n<p>constrained models</p>\n<p>Foundational framework</p>\n<p>The sloppy model framework emerged from computational biology and statistical physics,</p>\n<p>establishing that multiparameter models generically have parameter sensitivities spanning</p>\n<p>many orders of magnitude. Brown and Sethna (2003) introduced the concept, showing</p>\n<p>that biochemical signaling models with ~48 parameters have Hessian eigenvalue spectra</p>\n<p>spanning many decades, with only a few “stiff” parameter combinations controlling</p>\n<p>predictions and many “sloppy” directions allowing parameters to vary freely (Brown KS,</p>\n<p>Sethna JP, “Statistical mechanical approaches to models with many poorly known</p>\n<p>parameters,” Physical Review E 68, 021904, 2003). PubMed</p>\n<p>Waterfall et al. (2006) demonstrated that sloppiness constitutes a universality class by</p>\n<p>connecting the characteristic eigenvalue spectrum to Vandermonde matrix ensembles. The</p>\n<p>logarithms of eigenvalues are roughly uniformly spaced — a signature now recognized</p>\n<p>across physics, biology, and engineering. PubMed American Physical Society Crucially, this</p>\n<p>paper explicitly noted that sloppiness had also been demonstrated “in three multiparameter</p>\n<p>interatomic potentials fit to electronic structure” Cornell (Waterfall JJ, Casey FP,</p>\n<p>Gutenkunst RN, Brown KS, Myers CR, Brouwer PW, Elser V, Sethna JP, “Sloppy-model</p>\n<p>universality class and the Vandermonde matrix,” Physical Review Letters 97, 150601, 2006).</p>\n<p>Gutenkunst et al. (2007) tested 17 systems biology models and found that every one</p>\n<p>exhibited sloppy parameter sensitivities, cementing the universality claim. They argued that</p>\n<p>collective fits yield well-constrained predictions even when individual parameters remain</p>\n<p>poorly determined, and that direct parameter measurements must be “formidably precise</p>\n<p>and complete” to be more useful than collective fitting PubMed (Gutenkunst RN, Waterfall</p>\n<p>JJ, Casey FP, Brown KS, Myers CR, Sethna JP, “Universally sloppy parameter sensitivities in</p>\n<p>systems biology models,” PLoS Computational Biology 3(10), e189, 2007).</p>\n<p>Geometric and information-theoretic development</p>\n<p>The geometric interpretation of sloppiness matured through a series of papers by</p>\n<p>Transtrum, Machta, and Sethna. Transtrum et al. (2010) introduced the model manifold</p>\n<p>perspective: the set of all possible predictions as parameters vary forms a “hyper-ribbon” in</p>\n<p>data space with a geometric hierarchy of widths. This narrowness explains both why fitting</p>\n<p>is computationally difficult and why predictions are low-dimensional arXiv (Transtrum MK,</p>\n<p>Machta BB, Sethna JP, “Why are nonlinear fits to data so challenging?” Physical Review</p>\n<p>Letters 104, 060201, 2010). Transtrum et al. (2011) provided the comprehensive</p>\n<p>geometric treatment, showing that model manifolds universally exhibit a geometric series of</p>\n<p>widths, extrinsic curvatures, and parameter-effect curvatures American Physical Society</p>\n<p>(Transtrum MK, Machta BB, Sethna JP, “Geometry of nonlinear least squares with</p>\n<p>applications to sloppy models and optimization,” Physical Review E 83, 036701, 2011).</p>\n<p>Machta et al. (2013) connected sloppiness to fundamental physics by showing that</p>\n<p>parameter space compression underlies emergent theories. Using the Fisher Information</p>\n<p>Matrix (FIM) as a Riemannian metric on parameter space, they demonstrated that the</p>\n<p>emergence of effective theories — continuum limits, renormalization group fixed points —</p>\n<p>corresponds to compression from many microscopic parameters to few macroscopic ones.</p>\n<p>AIP Publishing PubMed Stiff FIM directions become emergent parameters; sloppy</p>\n<p>directions correspond to irrelevant microscopic detail arXiv (Machta BB, Chachra R,</p>\n<p>Transtrum MK, Sethna JP, “Parameter space compression underlies emergent theories and</p>\n<p>predictive models,” Science 342(6158), 604–607, 2013). This framework directly implies</p>\n<p>that the ~1.66 effective dimensions for elastic constant errors correspond to approximately</p>\n<p>two emergent combinations of potential parameters dominating elastic response.</p>\n<p>Transtrum et al. (2014) introduced the Manifold Boundary Approximation Method (MBAM)</p>\n<p>for systematic model reduction by following geodesics to the manifold boundary, removing</p>\n<p>sloppy parameter combinations while preserving stiff predictions PubMed (Transtrum MK,</p>\n<p>Qiu P, “Model reduction by manifold boundaries,” Physical Review Letters 113, 098701,</p>\n<p>2014). The comprehensive review by Transtrum et al. (2015) synthesized the full program,</p>\n<p>including FIM-based analysis, hyperribbon structure, MBAM, and connections to emergent</p>\n<p>theories across physics and biology AIP Publishing ADS (Transtrum MK, Machta BB, Brown</p>\n<p>KS, Daniels BC, Myers CR, Sethna JP, “Perspective: Sloppiness and emergent theories in</p>\n<p>physics, biology, and beyond,” Journal of Chemical Physics 143(1), 010901, 2015).</p>\n<p>Quinn et al. (2019) provided the first rigorous mathematical explanation for sloppiness,</p>\n<p>using Chebyshev approximation theory to derive universal bounds on model manifold widths</p>\n<p>as a consequence of model smoothness alone American Physical Society (Quinn KN, Wilber H,</p>\n<p>Townsend A, Sethna JP, “Chebyshev approximation and the global geometry of model</p>\n<p>predictions,” Physical Review Letters 122, 158302, 2019). The most recent comprehensive</p>\n<p>review is Quinn et al. (2023), covering model manifold structure, MBAM, information</p>\n<p>topology, optimal Bayesian priors, and visualization via intensive PCA PubMed Central</p>\n<p>(Quinn KN, Abbott MC, Transtrum MK, Machta BB, Sethna JP, “Information geometry for</p>\n<p>multiparameter models: New perspectives on the origin of simplicity,” Reports on Progress</p>\n<p>in Physics 86, 035901, 2023).</p>\n<p>Applications to interatomic potentials</p>\n<p>The seminal application of sloppy model theory to force fields is Frederiksen et al. (2004),</p>\n<p>which developed a Bayesian ensemble approach for estimating prediction errors from</p>\n<p>interatomic potentials. Working with EAM-type potentials for molybdenum fitted to DFT</p>\n<p>force databases, they showed that the potentials exhibit sloppy eigenvalue spectra identical</p>\n<p>in character to biological models, and that Bayesian error bars on elastic constants,</p>\n<p>gamma-surface energies, structural energies, and dislocation properties provided realistic</p>\n<p>uncertainty estimates. PubMed Parameters varied wildly across the ensemble while</p>\n<p>predictions remained constrained Cornell — the hallmark of low-dimensional prediction</p>\n<p>structure PLOS (Frederiksen SL, Jacobsen KW, Brown KS, Sethna JP, “Bayesian ensemble</p>\n<p>approach to error estimation of interatomic potentials,” Physical Review Letters 93, 165501,</p>\n<p>2004).</p>\n<p>Wen et al. (2017) applied Fisher information theory to a Stillinger-Weber potential for MoS₂,</p>\n<p>computing the FIM eigenvalue spectrum and verifying parameter identifiability. They used</p>\n<p>the geodesic Levenberg-Marquardt algorithm for fitting and demonstrated that the FIM</p>\n<p>analysis provides uncertainty bounds for elastic constants and other predicted properties</p>\n<p>(Wen M, Li J, Brommer P, Elliott RS, Sethna JP, Tadmor EB, “A force-matching Stillinger-</p>\n<p>Weber potential for MoS₂: Parameterization and Fisher information theory based sensitivity</p>\n<p>analysis,” Journal of Applied Physics 122, 244301, 2017).</p>\n<p>Kurniawan et al. (2022) provided the most thorough study of parametric uncertainty in</p>\n<p>classical potentials using the OpenKIM framework. Working with Lennard-Jones, Morse, and</p>\n<p>Stillinger-Weber potentials, they confirmed that interatomic potentials are sloppy models</p>\n<p>arXiv with bounded manifolds exhibiting a hierarchy of widths DNTB and low effective</p>\n<p>dimensionality. arXiv Many parameter combinations are unidentifiable, ResearchGate yet</p>\n<p>predictions remain constrained (Kurniawan Y, Petrie CL, Williams KJ, Transtrum MK, Tadmor</p>\n<p>EB, Elliott RS, Karls DS, Wen M, “Bayesian, frequentist, and information geometric</p>\n<p>approaches to parametric uncertainty quantification of classical empirical interatomic</p>\n<p>potentials,” Journal of Chemical Physics 156(21), 214103, 2022).</p>\n<p>Mortensen et al. (2005) extended the Bayesian sloppy-model framework to density</p>\n<p>functional theory itself, showing that DFT exchange-correlation functionals are also sloppy</p>\n<p>— sloppiness pervades the entire hierarchy of computational materials science (Mortensen</p>\n<p>JJ, Kaasbjerg K, Frederiksen SL, Nørskov JK, Sethna JP, Jacobsen KW, “Bayesian error</p>\n<p>estimation in density-functional theory,” Physical Review Letters 95, 216401, 2005).</p>\n<p>Neural networks and modern potentials</p>\n<p>Mao et al. (2024) showed that deep neural networks with widely different architectures,</p>\n<p>sizes, optimizers, and regularization all traverse the same low-dimensional manifold</p>\n<p>during training, with hyper-ribbon structure characteristic of sloppy models. ADS This</p>\n<p>suggests that neural network potentials (NequIP, MACE, etc.) should also produce low-</p>\n<p>dimensional elastic constant error patterns regardless of architecture (Mao J, Griniasty I,</p>\n<p>Teoh HK, Ramesh R, Yang R, Transtrum MK, Sethna JP, Chaudhari P, “The training process</p>\n<p>of many deep networks explores the same low-dimensional manifold,” Proceedings of the</p>\n<p>National Academy of Sciences 121(12), e2310002121, 2024). An analytical treatment</p>\n<p>followed in Mao et al. (2026), deriving conditions for hyper-ribbon emergence in linear</p>\n<p>models and characterizing phase boundaries (Mao J, Griniasty I, Sun Y, Transtrum MK,</p>\n<p>Sethna JP, Chaudhari P, “Analytical characterization of sloppiness in neural networks:</p>\n<p>Insights from linear models,” Physical Review E 113, 015306, 2026).</p>\n<p>Area 2: Simpson’s paradox and meta-analytic methodology for</p>\n<p>grouped correlations</p>\n<p>Foundational references</p>\n<p>Simpson (1951) demonstrated that collapsing two-way contingency tables can reverse</p>\n<p>associations, establishing the foundational paradox Wiley Online Library (Simpson EH, “The</p>\n<p>interpretation of interaction in contingency tables,” Journal of the Royal Statistical Society:</p>\n<p>Series B 13(2), 238–241, 1951). Blyth (1972) formalized the paradox and connected it to</p>\n<p>Savage’s sure-thing principle, Wikipedia providing the probabilistic framework (Blyth CR,</p>\n<p>“On Simpson’s paradox and the sure-thing principle,” Journal of the American Statistical</p>\n<p>Association 67(338), 364–366, 1972).</p>\n<p>The most celebrated real-world demonstration is the Berkeley admissions case (Bickel et</p>\n<p>al., 1975), where aggregate data showed apparent bias against women that reversed at the</p>\n<p>department level Science ResearchGate — women applied disproportionately to</p>\n<p>competitive departments, a confounding structure directly analogous to pooling elastic</p>\n<p>constant errors across elements with different physical properties (Bickel PJ, Hammel EA,</p>\n<p>O’Connell JW, “Sex bias in graduate admissions: Data from Berkeley,” Science 187(4175),</p>\n<p>398–404, 1975).</p>\n<p>Pearl (2014) argued that Simpson’s paradox can only be resolved through causal</p>\n<p>reasoning, introducing the do-calculus and back-door criterion for determining whether</p>\n<p>pooled or stratified analysis gives the correct inference. In the interatomic potential context,</p>\n<p>element identity acts as a causal confounder through distinct physical properties — atomic</p>\n<p>radius, electron configuration, bonding character — that independently influence both the</p>\n<p>potential’s error structure and the target elastic constants (Pearl J, “Comment:</p>\n<p>Understanding Simpson’s paradox,” The American Statistician 68(1), 8–13, 2014; also Pearl</p>\n<p>J, Causality: Models, Reasoning, and Inference, 2nd ed., Cambridge University Press,</p>\n<p>2009).</p>\n<p>Robinson (1950) established the closely related ecological fallacy, showing that state-level</p>\n<p>correlations between demographic variables can completely reverse from individual-level</p>\n<p>correlations. Wikipedia The pooled interatomic potential error correlation across elements</p>\n<p>is precisely this type of “ecological correlation” (Robinson WS, “Ecological correlations and</p>\n<p>the behavior of individuals,” American Sociological Review 15(3), 351–357, 1950).</p>\n<p>Correlation testing methodology</p>\n<p>Jackson and Somers (1991) is the key methodological reference for handling spurious</p>\n<p>correlations in grouped data. They demonstrated that ratios and indices sharing common</p>\n<p>components generate correlations even from random data, and advocated randomization</p>\n<p>(permutation) tests that generate a null distribution accounting for mathematical coupling.</p>\n<p>Their central recommendation is that the appropriate null correlation is frequently non-zero,</p>\n<p>and standard tests against r = 0 are inappropriate when variables share structural</p>\n<p>components PubMed Springer (Jackson DA, Somers KM, “The spectre of ‘spurious’</p>\n<p>correlations,” Oecologia 86(1), 147–151, 1991. DOI: 10.1007/BF00317404).</p>\n<p>Archie (1981) established the mathematical coupling problem: when derived variables</p>\n<p>share components (e.g., X/Z versus Y/Z), significant correlations arise even from</p>\n<p>uncorrelated raw data. This is the foundational reference for the non-zero null hypothesis</p>\n<p>issue — physical constraints among elastic constants (e.g., the Cauchy relation, stability</p>\n<p>requirements) create inherent baseline correlations that must not be attributed to</p>\n<p>systematic potential errors (Archie JP, “Mathematic coupling of data: A common source of</p>\n<p>error,” Annals of Surgery 193(3), 296–303, 1981).</p>\n<p>Kievit et al. (2013) provided a practical detection guide for Simpson’s paradox in</p>\n<p>continuous bivariate data with categorical grouping variables, including an R toolbox and</p>\n<p>statistical markers for the paradox. Frontiers Their methods are directly applicable to</p>\n<p>elastic constant error data grouped by element (Kievit RA, Frankenhuis WE, Waldorp LJ,</p>\n<p>Borsboom D, “Simpson’s paradox in psychological science: A practical guide,” Frontiers in</p>\n<p>Psychology 4, Article 513, 2013).</p>\n<p>Random-effects meta-analysis for correlations</p>\n<p>The standard methodology for properly combining correlations across heterogeneous</p>\n<p>groups uses Fisher’s z-transformation within a random-effects framework. Hedges and</p>\n<p>Olkin (1985) established the foundational approach: transform each group-specific</p>\n<p>correlation using z = arctanh(r), combine via inverse-variance weighting, and back-</p>\n<p>transform (Hedges LV, Olkin I, Statistical Methods for Meta-Analysis, Academic Press,</p>\n<p>1985). Borenstein et al. (2009) is the standard textbook reference, explaining fixed versus</p>\n<p>random-effects models, heterogeneity assessment (Q-statistic, I²</p>\n<p>, τ²), and interpretation</p>\n<p>(Borenstein M, Hedges LV, Higgins JPT, Rothstein HR, Introduction to Meta-Analysis, Wiley,</p>\n<p>2009). Borenstein et al. (2010) clarified when random-effects models are appropriate —</p>\n<p>specifically when group-specific effects (element-specific correlations) are expected to</p>\n<p>differ, as they do for different BCC metals (Borenstein M, Hedges LV, Higgins JPT, Rothstein</p>\n<p>HR, “A basic introduction to fixed-effect and random-effects models for meta-analysis,”</p>\n<p>Research Synthesis Methods 1(2), 97–111, 2010).</p>\n<p>Welz et al. (2022) proposed improved confidence intervals for the Fisher z approach in</p>\n<p>random-effects meta-analysis of correlations, showing that standard intervals can be</p>\n<p>unsatisfactory and providing enhanced variance estimators (Welz T, Viechtbauer W, Pauly</p>\n<p>M, “Fisher transformation based confidence intervals of correlations in fixed- and random-</p>\n<p>effects meta-analysis,” British Journal of Mathematical and Statistical Psychology 75(1), 1–</p>\n<p>38, 2022). DerSimonian and Laird (1986) provide the most widely used random-effects</p>\n<p>estimator for between-study variance (DerSimonian R, Laird N, “Meta-analysis in clinical</p>\n<p>trials,” Controlled Clinical Trials 7, 177–188, 1986). Higgins et al. (2003) introduced the I²</p>\n<p>statistic for quantifying heterogeneity (Higgins JPT, Thompson SG, Deeks JJ, Altman DG,</p>\n<p>“Measuring inconsistency in meta-analyses,” BMJ 327, 557–560, 2003).</p>\n<p>Simpson’s paradox in the physical sciences</p>\n<p>Examples of Simpson’s paradox in hard sciences remain rare, making instances in materials</p>\n<p>science especially noteworthy. Selvitella (2017) demonstrated the paradox in quantum</p>\n<p>mechanics (quantum harmonic oscillator, nonlinear Schrödinger equation) and in geometric</p>\n<p>and linear algebraic settings (Selvitella A, “The ubiquity of the Simpson’s paradox,” Journal</p>\n<p>of Statistical Distributions and Applications 4, Article 2, 2017). Chuang et al. (2009)</p>\n"}