{"id":"tda-error-landscapes","title":"Topological Data Analysis of Error Landscapes","subtitle":"Persistent homology for high-dimensional error surfaces.","category":"methods","tags":["tda","topology"],"source":"articles/docs/tda_error_landscapes_report.md","lang":"en","words":5853,"readMinutes":27,"toc":[{"depth":2,"text":"A Comprehensive Review of Persistent Homology Methods","id":"a-comprehensive-review-of-persistent-homology-methods"}],"html":"<h1 id=\"topological-data-analysis-of-interatomic-potential-error-landscapes\">Topological Data Analysis of Interatomic Potential Error Landscapes</h1><h2 id=\"a-comprehensive-review-of-persistent-homology-methods\">A Comprehensive Review of Persistent Homology Methods</h2><p><strong>Tags:</strong> Computational Physics, Machine Learning, Data Analysis</p>\n<hr>\n<p>Topological Data Analysis of Interatomic Potential Error Landscapes: A Comprehensive Review\nTopological Data Analysis of Interatomic Potential Error Landscapes: A Comprehensive Review</p>\n<ol>\n<li>Theoretical Foundations of TDA for Physical Systems\n1.1 Core Topological Concepts\n1.1.1 Persistent Homology: Filtrations, Barcodes, and Persistence Diagrams\nrepresents the cornerstone methodological framework for extracting robust, multiscale topological fe\nPersistent homology\nfiltrations\nFor interatomic potential error landscapes, this construction begins with a finite set of points in \nconnected components, loops, voids, and higher-dimensional cavities\nThe of each feature, quantified as the difference between its death and birth scale parameters, pro.\npersistence\npersistence barcodes\npersistence diagrams\nThe distance from the diagonal directly encodes feature significance, enabling quantitative comparis\nbottleneck distance\nWasserstein distance\nsublevelset filtration\nThe computational pipeline for persistent homology involves: (1) point cloud construction from raw d\nVietoris-Rips\nalpha complexes\n1.1.2 Betti Numbers as Topological Invariants Across Scales\nβₖ constitute the fundamental topological invariants that persistent homology tracks through the fil\nBetti numbers\ncounts\nconnected components\ncounts\none-dimensional loops or cycles\ncounts\ntwo-dimensional voids or cavities\nHigher βₖ extend to\nk-dimensional holes\nIn the context of persistent homology, Betti numbers become , producing or that reveal the multi-s..\nfunctions of the filtration parameter\nBetti curves\nBetti sequences\nFor materials science applications, Betti curves provide compact signatures of structural complexity\nBetti-0 curve\nBetti-1 and Betti-2 curves\nThe connection to error landscapes emerges when considering that interatomic potentials must accurat\nentire topological structure of configuration space\ntopological failures\nThe stability of Betti curves under perturbations—formalized through the stability theorems of persi\n1.1.3 Sublevelset Persistent Homology for Energy and Error Landscapes\noffers a natural construction for analyzing functions defined on point clouds or manifolds, making i\nSublevelset persistent homology\nf: X → ℝ\nXₐ = {x ∈ X : f(x) ≤ a}\nAs the threshold increases, these sublevelsets form a nested sequence , creating a filtration of the\nXₐ ⊆ X_b for a ≤ b\nFor potential energy landscapes, this construction has direct physical interpretation. At very low e\nWhen applied to —where the function represents |E_predicted − E_reference| or a similar discrepancy \nerror landscapes\nisolated high-error pockets may be kinetically inaccessible and thus benign, while percolating error\nThe mathematical foundation for sublevelset persistence rests on and its generalizations. For a Mor.\nMorse theory\n1.2 Historical Development of TDA in Computational Physics\n1.2.1 Early Applications to Molecular Configuration Spaces\nThe application of topological methods to molecular systems has roots extending back several decades\nKaczynski, Mischaikow, and Mrozek\nA landmark contribution came from the work of , who systematically developed persistent homology app\nXia and Wei\nThe key insight from these early applications was that . The number of persistent minima, the connec\nmolecular configuration spaces, despite their high nominal dimensionality (3N coordinates for N atom\nFor small molecule systems, persistent homology of potential energy landscapes enabled quantitative \nmany competing low-energy states and tortuous transition pathways\n1.2.2 Extension to Potential Energy Surfaces and Materials Pore Structures\nThe transition from molecular to extended materials systems required significant methodological adap\nporous materials and their characterization through topological analysis of pore structures\npresent a natural application domain because their function—molecular storage, separation, catalysis\nPorous materials\npore connectivity, cage sizes, and channel dimensionality\nLee et al. (2017)\nFor crystalline materials more broadly, the application of persistent homology required developing a\nperiodic boundary conditions\nVoronoi-based constructions\nThe extension to introduced additional complexity because the relevant configuration space includes.\npotential energy surfaces of materials systems\nlattice parameters, defect configurations, and potentially chemical composition variables\nsurprisingly complex topological structures\n1.2.3 Recent Advances in Multidimensional Persistence (2021–2026)\nThe period from 2021 to 2026 has witnessed substantial methodological advances that expand the appli\n, which tracks topological features across multiple filtration parameters simultaneously, has emerge\nMultidimensional persistent homology\nenergy versus force accuracy in potential fitting\nThe mathematical theory of multiparameter persistence has matured considerably during this period, w\nsinvariants\nrank invariant\ninterval decomposition approaches\ntemperature and pressure\nchemical composition and structural order parameter\nRIVET software\nA second major advance has been the development of methods that go beyond persistent homology to cap\nMerge trees and contour trees\nReeb graphs\nThe introduced in 2026 exemplifies these advances, implementing , with specific application to cata.\nLandscaper package\nmultidimensional TDA for probing loss landscapes in scientific machine learning\n1.3 Mathematical Framework for High-Dimensional Error Manifolds\n1.3.1 Point Cloud Construction from Simulation Data\nThe practical application of persistent homology to interatomic potential error landscapes begins wi\nsampling strategy, feature representation, and distance metric\nFor molecular dynamics simulations, the most direct point cloud construction samples configurations \n3N for N atoms in three dimensions\nexplicit dimensionality reduction\nimplicit low-dimensional structure<div class=\"table-wrap\"><table><caption>Table 1: Comparison of point cloud construction strategies for materials TDA</caption><thead><tr><th>Point Cloud Construction Approach</th><th>Advantages</th><th>Limitations</th><th>Typical Applications</th></tr></thead><tbody><tr><td data-label=\"Point Cloud Construction Approach\">Raw Cartesian coordinates</td><td data-label=\"Advantages\">Complete information, simple implementation</td><td data-label=\"Limitations\">High dimensionality, redundancy from translational/rotational invariance</td><td data-label=\"Typical Applications\">Small systems, validation studies</td></tr><tr><td data-label=\"Point Cloud Construction Approach\">Internal coordinates (bonds, angles, dihedrals)</td><td data-label=\"Advantages\">Natural invariance, physically interpretable</td><td data-label=\"Limitations\">Coordinate singularities, incomplete for large deformations</td><td data-label=\"Typical Applications\">Molecular conformational analysis</td></tr><tr><td data-label=\"Point Cloud Construction Approach\">SOAP/ACSF descriptors</td><td data-label=\"Advantages\">Rotationally invariant, chemically informative, fixed dimension</td><td data-label=\"Limitations\">Approximation of full configuration, hyperparameter sensitivity</td><td data-label=\"Typical Applications\">Machine learning potential training</td></tr><tr><td data-label=\"Point Cloud Construction Approach\">Neural network embeddings</td><td data-label=\"Advantages\">Learned relevance, adaptive compression</td><td data-label=\"Limitations\">Black-box nature, training data dependence</td><td data-label=\"Typical Applications\">Large-scale screening, transfer learning</td></tr><tr><td data-label=\"Point Cloud Construction Approach\">Voronoi-based constructions</td><td data-label=\"Advantages\">Natural periodicity handling, geometric interpretability</td><td data-label=\"Limitations\">Limited to moderate dimensions, computational cost</td><td data-label=\"Typical Applications\">Crystalline materials, defect analysis</td></tr></tbody></table></div></li>\n</ol>\n<p>of atomic trajectories provides one common reduction approach, projecting configurations onto the do\nPrincipal component analysis (PCA)\nnonlinear dimensionality reduction techniques\nFor systematic landscape exploration rather than dynamical sampling, more structured point cloud con\nGrid-based sampling\nsymmetry-restricted energy landscape approach\n1.3.2 Vietoris-Rips and Alpha Complex Filtrations\nGiven a point cloud representation, the construction of a filtration for persistent homology computa\nVietoris-Rips complex\nalpha complex</p>\n<div class=\"table-wrap\"><table><caption>Table 2: Comparison of simplicial complex filtrations for TDA</caption><thead><tr><th>Filtration Type</th><th>Definition</th><th>Computational Complexity</th><th>Best Suited For</th></tr></thead><tbody><tr><td data-label=\"Filtration Type\">Vietoris-Rips</td><td data-label=\"Definition\">Flag complex from pairwise distances ≤ 2ε</td><td data-label=\"Computational Complexity\">O(n³) worst case, often better with sparsity</td><td data-label=\"Best Suited For\">High dimensions, metric-only data, approximate analysis</td></tr><tr><td data-label=\"Filtration Type\">Alpha complex</td><td data-label=\"Definition\">Delaunay-restricted, empty circumsphere condition</td><td data-label=\"Computational Complexity\">O(n^(d/2)) in d dimensions, ~O(n²) for d≤3</td><td data-label=\"Best Suited For\">Low-dimensional embeddings, geometric accuracy</td></tr><tr><td data-label=\"Filtration Type\">Witness complex</td><td data-label=\"Definition\">Landmark-based approximation</td><td data-label=\"Computational Complexity\">O(kn) for k landmarks</td><td data-label=\"Best Suited For\">Very large datasets, exploratory analysis</td></tr><tr><td data-label=\"Filtration Type\">Graph-induced complex</td><td data-label=\"Definition\">Neighborhood graph-based</td><td data-label=\"Computational Complexity\">Near-linear with sparse graphs</td><td data-label=\"Best Suited For\">Network-structured data</td></tr></tbody></table></div>\n\n<p>The at scale ε includes all simplices whose vertices are pairwise within distance ε of each other. .\nVietoris-Rips complex\nflag complex\nRipser software\nThe offers an alternative that is particularly well-suited for Euclidean point clouds and provides .\nalpha complex\ngeometrically more accurate approximations\nhomotopy equivalent to the union of balls\nFor materials applications with , the alpha complex construction requires adaptation to handle the t\nperiodic boundary conditions\nPeriodic alpha shapes\nGUDHI\n1.3.3 Stability Theorems and Robustness Under Perturbation\nA foundational strength of persistent homology is its : small changes in input data produce correspo\nstability\nbottleneck stability theorem\nThe states that for two functions f, g: X → ℝ on a topological space, the bottleneck distance betwe.\nbottleneck stability theorem\nd_B(Dgm(f), Dgm(g)) ≤ ||f − g||_∞\nFor point cloud data, an analogous result bounds the persistence diagram perturbation in terms of th\nHausdorff distance\ntopological features with large persistence are robustly determined\nFor interatomic potential applications, this stability has several important implications:\n: Provided sampling is sufficiently dense that the Hausdorff distance to the true landscape is small\nJustification of finite sampling\n: The bottleneck or Wasserstein distance between persistence diagrams provides a that is itself sta.\nQuantitative comparison between potentials\nmetric of topological similarity\n: Stability ensures generalization from training to test data when topological features are used as \nSupport for persistence-based machine learning\nRecent extensions of stability theory address more nuanced aspects of robustness. The between persi.\ninterleaving distance\nprobabilistic stability results\n2. Sloppy Model Theory and Parameter Space Geometry\n2.1 Foundations of Sloppy Models in Interatomic Potentials\n2.1.1 Universally Sloppy Parameter Sensitivities in Physical Models\n, developed primarily by at Cornell University, provides a fundamental perspective on why complex p.\nSloppy model theory\nJames Sethna and collaborators\nsystems biology, chemical kinetics, and interatomic potentials\nexponential decay\nFor interatomic potentials, this sloppiness manifests in fitting landscapes where the cost function \nextended, nearly flat valleys along sloppy directions\nEmbedded Atom Method (EAM)\n17-parameter molybdenum potential\n10⁶ in magnitude\n5 stiff directions\nThe in interatomic potentials lies in the . Short-range repulsion is strongly constrained by core o.\nphysical origin of sloppiness\nhierarchical structure of atomic interactions</p>\n<div class=\"table-wrap\"><table><caption>Table 3: Key characteristics of sloppy models in interatomic potential fitting</caption><thead><tr><th>Sloppiness Characteristic</th><th>Manifestation in Interatomic Potentials</th><th>Implications</th></tr></thead><tbody><tr><td data-label=\"Sloppiness Characteristic\">Exponential eigenvalue decay</td><td data-label=\"Manifestation in Interatomic Potentials\">λₖ ~ exp(−αk/P) across 6–12 orders of magnitude</td><td data-label=\"Implications\">Effective dimensionality &lt;&lt; nominal parameter count</td></tr><tr><td data-label=\"Sloppiness Characteristic\">Long narrow canyons</td><td data-label=\"Manifestation in Interatomic Potentials\">Flat valleys in sloppy parameter directions</td><td data-label=\"Implications\">Slow optimization convergence, ensemble diversity</td></tr><tr><td data-label=\"Sloppiness Characteristic\">Parameter correlations</td><td data-label=\"Manifestation in Interatomic Potentials\">Compensating changes preserve predictions</td><td data-label=\"Implications\">Individual parameter uncertainties misleading</td></tr><tr><td data-label=\"Sloppiness Characteristic\">Prediction robustness</td><td data-label=\"Manifestation in Interatomic Potentials\">Stiff directions control observables</td><td data-label=\"Implications\">Reliable predictions despite parameter uncertainty</td></tr></tbody></table></div>\n\n<p>2.1.2 Eigenvalue Spectra of Cost Function Hessians: Exponential Decay Patterns\nThe quantitative characterization of sloppiness centers on the at the best-fit parameters. For a mo.\neigenvalue spectrum of the cost function Hessian\nP parameters\nN data points\nH_ij = ∂²C/∂θ_i∂θ_j\nλ₁ ≥ λ₂ ≥ ... ≥ λ_P\nλₖ ≈ λ₁ exp(−αk/P)\nfor some constant , indicating . This pattern implies that the , often exceeding , creating severe .\nexponential decay across the spectrum\ncondition number κ = λ₁/λ_P is enormous\n10⁶ for moderately complex models\nFor interatomic potentials, where , the eigenvalue spectrum provides essential diagnostic informatio\nP can range from tens (classical potentials) to millions (deep neural network potentials)\nmodel identifiability\nparameter combinations that are effectively unconstrained by the training data\nparameter combinations that must be precisely determined for accurate predictions\nSystematic studies across potential classes reveal :\ncharacteristic eigenvalue decay patterns</p>\n<div class=\"table-wrap\"><table><caption>Table 4: Sloppiness characteristics across interatomic potential classes</caption><thead><tr><th>Potential Class</th><th>Typical Parameter Count</th><th>Eigenvalue Span</th></tr></thead><tbody><tr><td data-label=\"Potential Class\">Effective Dimensionality</td><td data-label=\"Typical Parameter Count\">Lennard-Jones</td><td data-label=\"Eigenvalue Span\">10²–10³</td></tr><tr><td data-label=\"Potential Class\">Stillinger-Weber</td><td data-label=\"Typical Parameter Count\">7–11</td><td data-label=\"Eigenvalue Span\">10⁴–10⁶</td></tr><tr><td data-label=\"Potential Class\">~3–5</td><td data-label=\"Typical Parameter Count\">EAM/MEAM</td><td data-label=\"Eigenvalue Span\">10–50</td></tr><tr><td data-label=\"Potential Class\">10⁶–10⁸</td><td data-label=\"Typical Parameter Count\">~5–10</td><td data-label=\"Eigenvalue Span\">Neural network potentials</td></tr><tr><td data-label=\"Potential Class\">10⁴–10⁶</td><td data-label=\"Typical Parameter Count\">10¹⁰–10¹²</td><td data-label=\"Eigenvalue Span\">~10–100</td></tr></tbody></table></div>\n\n<p>2.1.3 Geometric Interpretation: Long Narrow Canyons and Flat Plateaus\nThe eigenvalue spectrum translates into a vivid . correspond to : moving perpendicular to these dir.\ngeometric picture of parameter space\nStiff directions\nsteep canyon walls\nSloppy directions\nlong, flat canyon floors\nhigh-dimensional river valley network\nThis geometry explains many empirical observations about potential fitting:\ndespite high nominal dimensionality reflects confinement to low-dimensional effective subspaces\nApparent success of simple optimization algorithms\n—cost plateaus above numerical precision—corresponds to exploration of increasingly flat sloppy dire\nDifficulty of achieving true convergence\nwith similar predictions but different parameters reflects the extensive degeneracy along sloppy dir\nExistence of many seemingly distinct &quot;good&quot; fits\nFor , the canyon picture must be augmented with considerations of . The cost landscape contains due.\nneural network potentials specifically\nnon-convexity\nmany local minima, saddle points, and possibly spurious valleys\nsloppy subspace\n&quot;effectively convex&quot; structure\n2.2 Connections Between Sloppiness and Topological Structure\n2.2.1 Parameter Space Geometry as Error Landscape Topology\nThe connection between sloppy model theory and topological data analysis arises from recognizing tha\nboth characterize high-dimensional landscapes\ntools\nlocal differential geometry\nglobal topology\nThe , central to sloppiness analysis, is fundamentally a : it describes curvature at a single point.\neigenvalue spectrum of the cost Hessian\nlocal quantity\nglobal organization of multiple minima, saddle points, and basins\nlocally sloppy (flat minima) yet globally complex (many minima)\nlocally stiff (sharp minima) yet globally simple (single basin)\ndirectly reveals . In a perfectly sloppy landscape with single minimum and flat sloppy directions, t\nPersistent homology of error landscapes\ntopological signatures of sloppiness\none infinite-persistence component\nmultiple competing minima, hierarchical basin structure\npersistent β₀ features\nbarrier heights relative to thermal or numerical noise\nThe of an error landscape provides a directly comparable to sloppiness eigenvalue spectra. Just as..\nmerge tree\nhierarchical summary\nmerge tree branches rank basins by barrier height\n&quot;topological eigenvalue spectrum&quot;\n2.2.2 Implications for Uncertainty Quantification in Potential Fitting\nTraditional uncertainty quantification in interatomic potentials , assuming . This fails dramaticall\npropagates parameter covariance to prediction variance\nlocal Gaussian approximations\nparameter covariance is enormous along sloppy directions\npredictions remain stable\n&quot;ensemble&quot; approach\n. The in parameter-space sublevelsets indicates whether ; reveal that any accepfit must sat...\nTopological analysis offers intermediate approaches\nnumber of connected components\ndistinct fitting solutions exist\npersistent features\nrobust constraints\npersistent error features\nall potentials in an ensemble agree\ndisagree\npersistence of these features\nFor practical potential development, these topological diagnostics suggest : rather than seeking a s\nensemble methods that deliberately explore the sloppy manifold\nmultiple optimization runs from diverse initializations sample distinct regions of the degenerate op\npredictive spread provides empirical uncertainty estimates\nrobust across the degenerate manifold\nsensitive to specific parameter choices\n2.2.3 Case Studies: Lennard-Jones and Stillinger-Weber Potentials\nThe and potentials provide concrete illustration of sloppiness-topology connections. The exhibits...\nLennard-Jones (LJ)\nStillinger-Weber (SW)\ntwo-parameter LJ potential (ε, σ)\nhighly correlated\ncompensating changes producing similar predictions\nequilibrium properties versus high-pressure behavior\ndifferent optimal parameters and different eigenvalue spectra\nvarying information content of different observations\nThe , with its , exhibits : in parameter space corresponding to ; and with some parameter combinat..\nStillinger-Weber potential for silicon\nthree-body angular terms\nricher structure\nmultiple local minima\ndifferent trade-offs between fitting cohesive energy, lattice constant, elastic constants, and defec\nhierarchical sloppiness\nTopological analysis of these landscapes\nglobal organization of these features\nnumber and distribution of local minima\nbarriers separating them\nconstraints that prevent simultaneous optimization of all target properties\nRecent have exploited sloppiness systematically. By , researchers have developed that retain accur..\nreparameterizations of SW using machine learning approaches\nidentifying the effective low-dimensional subspace of stiff directions\nreduced-parameter forms\nTopological validation\npreserve the essential basin structure\n2.3 Beyond Linear Analysis: Topological Features of Sloppy Landscapes\n2.3.1 Limitations of Hessian-Based Metrics and Spectral Density\nWhile provides valuable , it suffers from that motivate :\nHessian eigenvalue analysis\nlocal information about landscape geometry\nfundamental limitations\ntopological supplementation</p>\n<div class=\"table-wrap\"><table><caption>Table 5: Limitations of Hessian analysis and topological alternatives</caption><thead><tr><th>Limitation</th><th>Consequence for Potential Fitting</th><th>Topological Alternative</th></tr></thead><tbody><tr><td data-label=\"Limitation\">Purely local —describes curvature at single point</td><td data-label=\"Consequence for Potential Fitting\">Misses multiple optima, global connectivity</td><td data-label=\"Topological Alternative\">Persistent β₀ captures component structure</td></tr><tr><td data-label=\"Limitation\">Eigenvalue spectrum loses spatial structure</td><td data-label=\"Consequence for Potential Fitting\">Cannot distinguish single elongated valley from complex network</td><td data-label=\"Topological Alternative\">Merge trees encode hierarchical merging</td></tr><tr><td data-label=\"Limitation\">Assumes local quadratic behavior</td><td data-label=\"Consequence for Potential Fitting\">Fails at bifurcations, non-smooth points</td><td data-label=\"Topological Alternative\">Sublevelset persistence handles non-smoothness</td></tr><tr><td data-label=\"Limitation\">Ill-conditioned for sloppy directions</td><td data-label=\"Consequence for Potential Fitting\">Near-zero eigenvalues are numerical artifacts</td><td data-label=\"Topological Alternative\">Persistence provides robust significance filtering</td></tr><tr><td data-label=\"Limitation\">Insensitive to global topology</td><td data-label=\"Consequence for Potential Fitting\">Same spectrum for radically different landscapes</td><td data-label=\"Topological Alternative\">Betti curves distinguish topologically distinct cases</td></tr></tbody></table></div>\n\n<p>For , Hessian-based metrics can be . A minimum with appears favorable by sharpness-aware minimizati.\nerror landscapes specifically\nmisleading about generalization\nsmall largest eigenvalue (flat in all directions)\noverfitting\ninsufficient data constraint rather than genuine robustness\nsharp minimum with large eigenvalues\ntrue physical structure\neigenvalue magnitude alone cannot distinguish these cases\n2.3.2 Persistent Homology Capture of Non-Convex, Multi-Valley Structures\nThe for analyzing sloppy landscapes is its that Hessian analysis misses entirely. Each of the err...\ndefining capability of persistent homology\ndirect capture of non-convex, multi-valley structures\nlocal minimum\nbirth event in 0-dimensional persistent homology\npersistence of this component\nquantifies the depth and isolation of the minimum\nA landscape with thus exhibits , immediately flagging the that local analysis cannot detect. Simil..\nmultiple well-separated local minima\nmultiple long-persistence points in its 0-dimensional persistence diagram\nmultimodal structure\n1-dimensional homology captures loops\npaths that return to similar error states\ntrade-offs or conservation laws\npersistence of such loops\nrobustness of these redundancies\nFor , these topological features have :\ninteratomic potential fitting\ndirect physical interpretation\n→ (e.g., emphasizing bulk properties vs. surfaces)\nMultiple minima\ndistinct fitting strategies\n→ (e.g., two-body vs. three-body term compensation)\nLoops\ncyclic parameter trade-offs\n→ (e.g., simultaneous satisfaction of mutually incompatible constraints)\nHigher-dimensional voids\ninaccessible optimal combinations\n2.3.3 Topological Signatures of Model Stiffness and Parameter Correlations\nThe that complements traditional eigenvalue analysis. In models with —corresponding to sloppy direc.\ntopological structure of error landscapes encodes information about model stiffness and parameter co\nstrongly correlated parameters\nlower-dimensional structure\neffective dimensionality of low-cost regions is reduced\nsuppressed higher Betti numbers\n&quot;topological dimension reduction&quot;\nThe encodes :\ndistribution of persistence across filtration scales\nhierarchical structure in parameter constraints\nFeatures persisting over large scale ranges\nfundamental, physically-required constraints\nFeatures appearing only at intermediate scales\nfitting-data-dependent structure\nShort-persistence features\nnoise or overfitting\nFor , the , with that , then . These dynamics inform .\nneural network potentials\ntopological signature evolves during training\ninitial random parameterization showing simple topology\ncomplexifies as the network discovers effective representations\nsimplifies as the model settles into a robust basin\ntraining strategies and early stopping criteria\n3. Persistent Homology vs. Linear Methods: Revealing Hidden Structure\n3.1 PCA Eigenvalue Spectra and Their Limitations\n3.1.1 Standard Principal Component Analysis of Error Data\nremains the for analyzing high-dimensional simulation data, including interatomic potential error l.\nPrincipal Component Analysis (PCA)\ndominant dimensionality reduction technique\ndirections of maximum variance\northogonal basis ordered by importance\neigenvalue spectrum\neffective dimensionality\nFor well-behaved error distributions, PCA can be highly effective: a , enabling visualization and ef\nsmall number of principal components may capture the majority of variance\nfundamental assumption of PCA—that important structure aligns with directions of maximum variance—fa\ncurved, nonlinear manifolds where variance is distributed across many weak directions\nlinear subspaces, distorting their geometry and obscuring their topology\nThe mathematical limitations are clear: PCA seeks the , which is but . The eigenvalue spectrum may .\nbest linear approximation in mean-square sense\noptimal for Gaussian distributions and quadratic cost functions\nsuboptimal for complex, multi-modal landscapes\ngradual decay without clear cutoff\napparent structure (gaps, plateaus) that reflects sampling artifacts rather than genuine features\nsloppy models specifically\nmirrors the Hessian eigenvalue decay\nsimilarity obscures important differences in what the methods capture\n3.1.2 Linear Projection Artifacts in High-Dimensional Landscapes\nWhen PCA is applied to , the resulting linear projections introduce :\nnonlinear, curved manifolds\nsystematic artifacts that can mislead interpretation</p>\n<div class=\"table-wrap\"><table><caption>Table 6: Linear projection artifacts in PCA of error landscapes</caption><thead><tr><th>Artifact Type</th><th>Mechanism</th><th>Consequence for Error Analysis</th></tr></thead><tbody><tr><td data-label=\"Artifact Type\">Curvature distortion</td><td data-label=\"Mechanism\">Geodesic distances underestimated</td><td data-label=\"Consequence for Error Analysis\">Nearby points on manifold appear distant; distant points appear nearby</td></tr><tr><td data-label=\"Artifact Type\">Global structure loss</td><td data-label=\"Mechanism\">Separate components project together; single components split</td><td data-label=\"Consequence for Error Analysis\">False connectivity or fragmentation</td></tr><tr><td data-label=\"Artifact Type\">Scale ambiguity</td><td data-label=\"Mechanism\">Features at different scales compete for variance</td><td data-label=\"Consequence for Error Analysis\">Large-scale structure dominates; fine-scale topology obscured</td></tr><tr><td data-label=\"Artifact Type\">Concentration of measure</td><td data-label=\"Mechanism\">High-dimensional projections tend to Gaussianity</td><td data-label=\"Consequence for Error Analysis\">Distinctive structure erased</td></tr></tbody></table></div>\n\n<p>For , these artifacts have . Two potentials with —same number and connectivity of low-error regions—\ninteratomic potential comparison\nserious consequences\ntopologically equivalent error landscapes\nappear different under PCA\nfundamentally different topology\nproject to similar PCA representations\nuse of nonlinear, topology-preserving methods such as persistent homology\n3.1.3 Insensitivity to Non-Isotropic, Curved Manifold Structures\nA for analyzing error landscapes is its —precisely the structures that arise from . When data lies .\ncritical limitation of PCA\ninsensitivity to non-isotropic, curved manifold structures\nsloppy parameter correlations and complex model nonlinearities\ncurved submanifold\nmany components to achieve reasonable approximation\napparent high dimensionality reflecting curvature rather than true complexity\nFor , this insensitivity can be . Suppose errors vary primarily along a —error is low in reactant an\nerror landscapes on curved manifolds\ncatastrophic\ncurved reaction coordinate\ncapture configuration variance\nrather than error variation\ncompletely missing the chemically relevant structure\northogonal to the dominant variance directions\n. Because it depends —whether points are within distance ε, not their absolute positions—it . Curved\nPersistent homology, by contrast, is naturally insensitive to manifold curvature\nonly on proximity structure\nrespects the intrinsic metric of whatever manifold the data inhabits\nno differently from flat ones\npairwise distance matrix matters\ncoordinate-invariance\nmajor advantage for molecular systems where Cartesian coordinates are highly redundant and physicall\n3.2 Topological Features Beyond PCA Capture\n3.2.1 Detection of Loops, Voids, and Higher-Dimensional Holes\nThe is its in data manifolds—features that . These features have in error landscapes:\nmost immediate advantage of persistent homology over PCA\nsystematic detection and quantification of loops, voids, and higher-dimensional holes\nlinear methods fundamentally cannot capture\ndirect physical interpretations</p>\n<div class=\"table-wrap\"><table><caption>Table 7: Topological features and their physical interpretations</caption><thead><tr><th>Topological Feature</th><th>Physical Interpretation in Error Landscapes</th><th>Diagnostic Value</th></tr></thead><tbody><tr><td data-label=\"Topological Feature\">β₀: Connected components</td><td data-label=\"Physical Interpretation in Error Landscapes\">Distinct error regimes, isolated accurate regions</td><td data-label=\"Diagnostic Value\">Fragmentation vs. connectivity of good predictions</td></tr><tr><td data-label=\"Topological Feature\">β₁: Loops</td><td data-label=\"Physical Interpretation in Error Landscapes\">Cyclic error correlations, trade-off cycles, periodic constraints</td><td data-label=\"Diagnostic Value\">Systematic compensating errors, symmetry violations</td></tr><tr><td data-label=\"Topological Feature\">β₂: Voids</td><td data-label=\"Physical Interpretation in Error Landscapes\">\\&quot;Blind spots\\&quot; surrounded by accurate regions, inaccessible optima</td><td data-label=\"Diagnostic Value\">Training data gaps, extrapolation hazards</td></tr><tr><td data-label=\"Topological Feature\">β₃+ : Higher cavities</td><td data-label=\"Physical Interpretation in Error Landscapes\">Complex multi-parameter constraints</td><td data-label=\"Diagnostic Value\">Rare, but indicate severe model limitations</td></tr></tbody></table></div>\n\n<p>For , loops commonly arise from : rotating a crystal by 360° returns to the original configuration, \nmaterials systems\nperiodic boundary conditions and crystal symmetries\nloop in orientation space that PCA projects to a point\nconstraint satisfaction problems\nfeasible region may have complex topology with inaccessible interior regions\nThe —the —indicates their : features with , while . For , comparing persistence diagrams between pot\npersistence of these features\nscale range over which they exist\nrobustness\nlarge persistence represent fundamental structural constraints\nshort-lived features may be sampling artifacts\nmodel comparison\nwhich topological features are intrinsic to the physical system (common across potentials) versus ar\n3.2.2 Betti Number Evolution as Function of Filtration Scale\nThe ——provides a that . For error landscapes, this evolution reveals:\nevolution of Betti numbers as functions of filtration scale\nBetti curves βₖ(ε)\nrich, quantitative characterization of landscape structure\nstatic eigenvalue spectra cannot match</p>\n<div class=\"table-wrap\"><table><caption>Table 8: Betti curve interpretation for error landscape analysis</caption><thead><tr><th>Betti Curve Feature</th><th>Interpretation</th><th>Diagnostic Application</th></tr></thead><tbody><tr><td data-label=\"Betti Curve Feature\">Rapid β₀ decay</td><td data-label=\"Interpretation\">Efficient connectivity, single dominant basin</td><td data-label=\"Diagnostic Application\">Well-converged, robust potential</td></tr><tr><td data-label=\"Betti Curve Feature\">Extended β₀ plateau</td><td data-label=\"Interpretation\">Fragmented structure, multiple distinct regions</td><td data-label=\"Diagnostic Application\">Multi-modal errors, transferability concerns</td></tr><tr><td data-label=\"Betti Curve Feature\">β₁ peak position/height</td><td data-label=\"Interpretation\">Characteristic loop scale and abundance</td><td data-label=\"Diagnostic Application\">Cyclic constraint strength</td></tr><tr><td data-label=\"Betti Curve Feature\">β₂ emergence</td><td data-label=\"Interpretation\">Three-dimensional void structure</td><td data-label=\"Diagnostic Application\">Complex failure modes, data gaps</td></tr><tr><td data-label=\"Betti Curve Feature\">β₁/β₀ ratio</td><td data-label=\"Interpretation\">Landscape complexity vs. connectivity</td><td data-label=\"Diagnostic Application\">Ruggedness indicator</td></tr></tbody></table></div>\n\n<p>enables . The over ε measures ; the emphasizes . For the , these metrics reveal : show ; indicat...\nComparing Betti curves across potentials\nquantitative model comparison\nintegral of β₀(ε)\ntotal &quot;basin complexity&quot;\npersistence-weighted integral\nrobust features\n23-potential benchmark dataset\nclustering\nclassical potentials (EAM, MEAM)\nsimple signatures with rapid β₀ decay\nneural network potentials exhibit extended plateaus\nmany persistent minima\nGaussian approximation potentials show intermediate complexity\n3.2.3 Persistence Diagrams vs. Eigenvalue Histograms: Complementary Information\nabout landscape structure:\nPersistence diagrams and eigenvalue histograms provide complementary, non-redundant information</p>\n<div class=\"table-wrap\"><table><caption>Table 9: Complementary information from PCA and persistent homology</caption><thead><tr><th>Aspect</th><th>PCA Eigenvalue Histogram</th><th>Persistence Diagram</th></tr></thead><tbody><tr><td data-label=\"Aspect\">Information type</td><td data-label=\"PCA Eigenvalue Histogram\">Variance, linear correlation</td><td data-label=\"Persistence Diagram\">Topology, connectivity, shape</td></tr><tr><td data-label=\"Aspect\">Scale treatment</td><td data-label=\"PCA Eigenvalue Histogram\">Global, single-scale</td><td data-label=\"Persistence Diagram\">Multiscale, local-to-global</td></tr><tr><td data-label=\"Aspect\">Stability</td><td data-label=\"PCA Eigenvalue Histogram\">Sensitive to outliers</td><td data-label=\"Persistence Diagram\">Sunder perturbations</td></tr><tr><td data-label=\"Aspect\">Computational cost</td><td data-label=\"PCA Eigenvalue Histogram\">O(min(n²m, nm²))</td><td data-label=\"Persistence Diagram\">O(n³) worst case, often better</td></tr><tr><td data-label=\"Aspect\">Interpretability</td><td data-label=\"PCA Eigenvalue Histogram\">Familiar, widely used</td><td data-label=\"Persistence Diagram\">Requires topological training</td></tr><tr><td data-label=\"Aspect\">Missing structure</td><td data-label=\"PCA Eigenvalue Histogram\">Loops, voids, global connectivity</td><td data-label=\"Persistence Diagram\">Absolute scale, linear trends</td></tr></tbody></table></div>\n\n<p>For , the . should be computed to ; should be computed to ; and the —whether , whether —should be ..\npractical analysis\ncombination of both perspectives is often most powerful\nPCA eigenvalue spectra\ncharacterize effective dimensionality and stiff/sloppy decomposition\npersistence diagrams\ncharacterize basin structure and model robustness\nrelationship between these analyses\nstiff directions correspond to persistent features\nsloppy directions contain hidden topological structure\nexplicitly investigated\n3.3 Recent Methodological Developments (2021–2026)\n3.3.1 Multidimensional Persistent Homology for Loss Landscapes\nThe has seen , with in potential development. Traditional persistent homology tracks a ; enables ...\nextension from single-parameter to multiparameter persistent homology\nsubstantial methodological development from 2021–2026\ndirect applications to multi-objective loss landscapes\nsingle scale parameter\nmultidimensional persistence\nsimultaneous analysis of multiple error components\nmultiple physical conditions\nThe involves rather than , substantially complicating both . Key advances include:\nmathematical framework\nmodules over partially ordered sets (posets)\nvector spaces over the real line\ntheory and computation\n, enabling efficient computation\nDiscrete Morse theory for multifiltrations\nthat reduce complex modules to essential generators\nMinimal presentation algorithms\nincluding and that provide tracsummaries\nInvariant definitions\nrank invariants\nfibered barcodes\nFor , without arbitrary scalarization. Rather than combining energy and force errors into a single .\ninteratomic potential benchmarking\nmultidimensional persistence enables rigorous comparison across multiple performance criteria\nanalyzes the Pareto front structure directly\ntrade-offs and identifying potentials that are genuinely non-dominated\ntopological structure of Pareto fronts\nconnectedness, loops, voids\nimprovement in one criterion necessarily degrades others\nsynergistic improvement is possible\n3.3.2 Merge Trees and Contour Trees for Gradient-Based Analysis\nprovide that capture while relative to full persistent homology. The tracks , encoding the of l...\nMerge trees and contour trees\nsimplified topological representations\nhierarchical structure\nreducing computational complexity\nmerge tree\ncomponent evolution in sublevelsets\n&quot;basin hierarchy&quot;\ncontour tree\ntrack both sublevel and superlevel merging\ncomplete Reeb graph structure for simply connected domains\nFor , merge trees enable without explicit saddle point search. The : given two minima, their ; the .\npotential energy and error landscapes\nefficient computation of barrier heights, transition states, and minimum energy paths\ntree structure supports rapid queries\nlowest common ancestor in the merge tree gives the highest saddle point on the connecting path\npersistence of this merge gives the barrier height\nmerge trees in near-linear time\ninteractive exploration of large landscapes\nThe —using , or —enables . For , (number of minima, distribution of persistences, tree balance meas.\nintegration of merge trees with machine learning\ntree structure as input to graph neural networks\ntraining models to predict merge tree statistics from descriptors\nrapid landscape characterization without explicit simulation\npotential comparison\nmerge tree statistics\ncompact fingerprints that correlate with dynamical properties\nglass-forming ability or phase transition behavior\n3.3.3 Normalized Persistence Curves for Quantitative Model Comparison\n——enable of landscape structure. (by ) enable .\nPersistence curves\nparametric plots of Betti numbers or related quantities versus filtration scale\nquantitative comparison and statistical analysis\nNormalization strategies\ntotal persistence, maximum persistence, or integral characteristics\nmeaningful comparison across datasets of different size and scale\nFor , that facilitate of different models. Curves computed from , applied to , reveal : potentials..\ninteratomic potential benchmarking\nnormalized persistence curves provide standardized topological signatures\nsystematic ranking and classification\nerror landscapes of different potentials\ncommon validation sets\nsystematic differences in landscape complexity\nextended β₀ persistence\nfragmented error structure\nenhanced β₁ persistence\ncomplex compensating parameter relationships\nMachine learning on these curve features\nfunctional data or extracting scalar summaries\nautomated potential classification and performance prediction\nRecent has established for appropriately normalized curves, supporting their use in . enable , wi...\ntheoretical work\nstability properties and statistical convergence\nrigorous statistical comparison\nPersistence curve kernels\nsimilarity assessment and statistical testing\ncorrections for multiple comparisons across filtration scales\nrigorous comparison protocols\nessential for systematic potential assessment\n4. Applications to Energy and Error Landscapes in Materials Science\n4.1 Persistent Homology of Potential Energy Landscapes\n4.1.1 Sublevelset Analysis of n-Alkane Conformational Spaces\nThe represents a where persistent homology reveals . Despite their chemical simplicity, n-alkanes ..\nconformational analysis of n-alkanes\nparadigmatic application\nstructure invisible to traditional methods\ncomplex conformational landscapes governed by torsional degrees of freedom\nnumber of local minima growing exponentially with chain length\nconnectivity of transition pathways determining kinetic properties\nEarly applications demonstrated that . For , where complete enumeration is feasible, persistent homo\nsublevelset persistence of torsional energy landscapes captures the hierarchical organization of con\nbutane and pentane\ncorrectly identifies the number and persistence of distinct rotameric states\nβ₀ peaks corresponding to sconformations\nβ₁ features indicating the cyclic connectivity of torsional pathways\npersistence of these features correlates with experimental barriers measured by NMR spectroscopy\nFor , persistent homology of —whether from —provides . Studies of revealed that the when measured ..\nlonger chains where exhaustive enumeration becomes impossible\nsampled landscapes\nmolecular dynamics trajectories or specialized sampling algorithms\nstatistical characterization of landscape complexity\nC₁₀–C₂₀ alkanes\nnumber of persistent minima scales sub-exponentially with chain length\nphysically relevant persistence thresholds\nkinetic accessibility constraints effectively reduce the relevant state space\nBetti curve structure indicates increasing topological complexity with length\nself-similar organization that suggests hierarchical folding mechanisms\nThe arises naturally when . Classical force fields () make . of their predicted landscapes against..\nextension to error landscapes\ncomparing different force field representations of alkane energetics\nOPLS, CHARMM, AMBER\ndifferent approximations to torsional potentials, nonbonded interactions, and coupling terms\nPersistent homology comparison\nhigh-level quantum chemical reference\nsystematic topological differences\nover-stabilize extended conformations\nexcess persistent minima in β₀\nmisrepresent torsional coupling\nβ₁ structure of transition pathways\ntopological discrepancies\ninvisible to traditional RMSD or energy error metrics\ncorrelate with observed differences in predicted thermodynamic and kinetic properties\n4.1.2 Geometric Landscapes for Material Discovery: Pore Structure Characterization\nThe , introduced by , represents a : replacing . For , where , geometric landscapes use that capt...\ngeometric landscape framework\nKrishnapriyan et al. (2020)\nparadigm shift in materials informatics\nenergy-based structure search with topology-based similarity assessment\nporous molecular crystals\nenergy-structure-function (ESF) mapping requires evaluation of thousands to millions of candidate st\npersistent homology to define similarity metrics\npore connectivity and accessibility more effectively than conventional geometric descriptors\nThe construction proceeds by , computing , and defining . The resulting —structures arranged by —rev\nrepresenting each porous structure as a point cloud of pore centers or void space samples\npersistence diagrams that characterize pore topology\nlandscape distance through optimal transport metrics on these diagrams\n&quot;geometric landscape&quot;\ntopological similarity rather than energy\nstructure-property relationships obscured in energy-based projections\ngeometric landscape features achieved remarkable accuracy\nmean absolute error of 7.0 (v STP/v) and Spearman rank correlation of 0.95 for methane deliverable c\n600 training points\nThis success motivates . Just as , could reveal : , , whether . The , with .\nextension to error landscape analysis\ngeometric landscapes reveal structure-function relationships\n&quot;error landscapes&quot; in topological space\nstructure-error relationships\nwhich structural features consistently challenge particular potentials\nhow error correlates with topological complexity\ncertain persistence signatures predict potential failure\nmathematical framework transfers directly\npersistence diagrams of predicted versus reference configurations serving as feature descriptors\n4.1.3 Voronoi-Based Point Cloud Construction for Crystalline Materials\nrequire that . , building on , provide that .\nCrystalline materials\nspecialized point cloud constructions\nrespect periodic boundary conditions and capture the relevant structural features\nVoronoi-based methods\nestablished techniques in crystallography and computational geometry\nnatural frameworks\nconnect to physical understanding of atomic environments and their topological organization\nThe of a , producing a . The provides a . As the , simplices are added based on , producing a tha...\nVoronoi decomposition\nperiodic point set partitions space into regions closer to each point than to any other\ncell complex that encodes local coordination geometry\ndual Delaunay triangulation\nsimplicial complex that serves as the basis for alpha complex filtrations\nalpha parameter increases\nempty circumsphere conditions\nnested sequence of complexes\nhow atomic neighborhoods connect as the effective interaction range increases\nFor , the . produce that enable . More significantly, the : , while .\nelemental crystals and simple compounds\nalpha complex persistence reveals the topological signatures of different structure types\nFace-centered cubic, body-centered cubic, and hexagonal close-packed arrangements\ncharacteristically different barcode patterns\nautomatic structure identification and classification\npersistence of features indicates structural stability\nrobust structures maintain their topological signature across wide parameter ranges\nfragile structures show rapid feature evolution indicating sensitivity to perturbation\n4.2 Model Comparison and Benchmarking Methodologies\n4.2.1 Symmetry-Restricted Energy Landscapes for ML Potential Evaluation\nThe , developed by , provides a that . Rather than relying , this approach where , enabling .\nsymmetry-restricted energy landscape methodology\nParackal, Armiento, and Trybel (2026)\nsystematic framework for evaluating machine learning interatomic potentials\ndirectly addresses the limitations of traditional benchmarking approaches\nsolely on validation set errors or materials discovery metrics\nconstructs explicit two-dimensional slices through the potential energy surface\natomic positions are varied along selected Wyckoff degrees of freedom within fixed crystal symmetry\ndirect visual and quantitative comparison between predicted and reference landscapes\nThe lies in the : by , the that . The , which specify , provide for this reduction. Varying these...\nmethodological innovation\nconstrained exploration\npreserving crystal symmetry during deformation\nhigh-dimensional configuration space is reduced to manageable low-dimensional subspaces\nretain physical relevance\nWyckoff positions\nsymmetrically equivalent atomic sites in crystallographic space groups\nnatural collective coordinates\nsymmetry-preserving deformations\nlocal curvature of the energy landscape around equilibrium structures\nelastic response, phonon frequencies, and barrier heights for symmetry-allowed transitions\nFor , this methodology . Foundation models such as —universal pre-trained potentials intended for br\nmachine-learned potentials\nreveals artifacts that global error metrics may obscure\nMACE, CHGNet, M3GNet, ORB, and SevenNet\nvarying performance on these restricted landscapes\naccurately capture local minima and their curvatures but fail at larger displacements\nunphysical oscillations or incorrect barrier heights in specific symmetry channels\ninadequate representation of certain local environments despite good average performance\nThe extends the . reveals the , the , and the presence of any . provides a . Models with may sti...\ntopological analysis of these restricted landscapes\nvisual comparison to quantitative characterization\nPersistent homology of the two-dimensional energy surfaces\nnumber and persistence of local minima\nconnectivity of basins through saddle points\nspurious topological features\nComparison between predicted and DFT reference landscapes using bottleneck distances between persist\nsingle numerical score that captures topological fidelity\nsmall persistence distances\nquantitative energy errors\npreserve the correct landscape organization\nlarge persistence distances\nqualitatively incorrect topology\nunreliable dynamics regardless of energy accuracy\n4.2.2 Visual Comparison of PES Slices: Maxima, Minima, and Saddle Points\nThe , enabled by , provides that . ——that , making .\nvisual comparison of potential energy surface slices\ndimensionality reduction to two-dimensional subspaces\nintuitive assessment of model performance\ncomplements quantitative topological analysis\nHuman pattern recognition excels at identifying qualitative discrepancies\nmissing features, spurious oscillations, incorrect curvature\nautomated metrics may miss\nvisual inspection a valuable component of comprehensive benchmarking protocols\nEffective to . The , but : where harmonic approximations apply; ; (tensile, shear, bending); and ...\nvisual comparison requires careful selection of the two-dimensional subspace\nmaximize information content\nsymmetry-restricted approach ensures physical relevance\nadditional criteria guide specific slice selection\nproximity to equilibrium structures\ninclusion of known transition states or reaction pathways\ncoverage of chemically relevant deformation modes\nrepresentation of diverse local coordination environments\nFor , , motivating . However, ——remains that .\nlarge-scale benchmarking with ~12,000 materials and 23 potentials\nsystematic visual comparison is infeasible\nautomated topological summaries\nstrategic visual inspection of representative cases\nparticularly those where topological metrics indicate anomalies\nessential for validating automated assessments and identifying failure modes\nquantitative measures may not capture\n4.2.3 Large-Scale Benchmarking: 12,000-Material Datasets Across 23 Potentials\nThe ——presents . The is ; the is .\nscale of modern MLIP benchmarking\nthousands of materials, dozens of potentials\nboth opportunity and challenge for topological analysis\nopportunity\ncomprehensive characterization of model behavior across chemical and structural space\nchallenge\ncomputational cost and result interpretation\nFor a , :\nhypothetical 12,000-material dataset with 23 interatomic potentials\nmultiple analysis frameworks are possible</p>\n<div class=\"table-wrap\"><table><caption>Table 10: Analysis frameworks for large-scale potential benchmarking</caption><thead><tr><th>Analysis Framework</th><th>Point Cloud Construction</th><th>Computational Cost</th><th>Information Content</th></tr></thead><tbody><tr><td data-label=\"Analysis Framework\">Configuration-space</td><td data-label=\"Point Cloud Construction\">~12,000 points per material, 276,000 total, 23D error vectors</td><td data-label=\"Computational Cost\">Very high: O(n³) with n~10⁵–10⁶</td><td data-label=\"Information Content\">Complete error topology, infeasible</td></tr><tr><td data-label=\"Analysis Framework\">Error-space</td><td data-label=\"Point Cloud Construction\">23D points per configuration, 12,000 × N_config total</td><td data-label=\"Computational Cost\">Moderate: O(n³) with n~10⁴–10⁵</td><td data-label=\"Information Content\">Potential clustering, error correlations</td></tr><tr><td data-label=\"Analysis Framework\">Descriptor-space</td><td data-label=\"Point Cloud Construction\">SOAP/ACE embeddings, fixed dimension</td><td data-label=\"Computational Cost\">Moderate: O(n³) with n~10⁴–10⁵, d~100</td><td data-label=\"Information Content\">Physical interpretability, approximation</td></tr><tr><td data-label=\"Analysis Framework\">Stratified sampling</td><td data-label=\"Point Cloud Construction\">Representative subsets by chemistry/structure</td><td data-label=\"Computational Cost\">Low: O(n³) with n~10³–10⁴</td><td data-label=\"Information Content\">Scalable, may miss rare features</td></tr></tbody></table></div>\n\n<p>: of the full dataset, . :\nEstimated costs for full persistent homology\n~10⁶ CPU-hours for complete analysis\nprohibitive for routine evaluation\nApproximation strategies achieve practical costs</p>\n<div class=\"table-wrap\"><table><caption>Table 11: Approximation strategies for large-scale TDA</caption><thead><tr><th>Strategy</th><th>Speedup</th><th>Accuracy Trade-off</th><th>Application</th></tr></thead><tbody><tr><td data-label=\"Strategy\">Random subsampling (1%)</td><td data-label=\"Speedup\">10⁴×</td><td data-label=\"Accuracy Trade-off\">Misses rare features</td><td data-label=\"Application\">Initial screening</td></tr><tr><td data-label=\"Strategy\">Grid-based filtration</td><td data-label=\"Speedup\">10²×</td><td data-label=\"Accuracy Trade-off\">Discretization artifacts</td><td data-label=\"Application\">Regular structures</td></tr><tr><td data-label=\"Strategy\">Persistence thresholding</td><td data-label=\"Speedup\">10×</td><td data-label=\"Accuracy Trade-off\">Loss of weak features</td><td data-label=\"Application\">Final validation</td></tr><tr><td data-label=\"Strategy\">GPU acceleration (Ripser++)</td><td data-label=\"Speedup\">10–100×</td><td data-label=\"Accuracy Trade-off\">Hardware dependence</td><td data-label=\"Application\">Large-scale production</td></tr></tbody></table></div>\n\n<p>For the , a : , followed by . This .\n12,000-material, 23-potential benchmark\nhybrid approach achieves practical analysis\ngrid-based initial screening identifies materials with anomalous topology\nfull computation on selected systems\ntargets computational effort where topological differences are most informative\n4.3 Materials System-Specific Applications\n4.3.1 Metals: Defect Analysis and Dislocation Networks\npresent due to the **importance of defects—dislocations, grain boundaries, point defects\nMetallic systems\nparticular challenges for topological analysis\nInteractive Report\nGenerated, click to preview\nPreview\nReference\nStarting Deep Research creates a new chat.\ntextbox\nAsk Kimi to get an in-depth research report\n91 left\nNew chat\nHTML\nCode\nPreview\nSwitch Preview Mode\nRefresh\nShare\nGenerated by Kimi AI\nClaude is active in this tab group\nOpen chat\nDismiss\nViewport: 1440x788\nTab Context:</p>\n<ul>\n<li>Executed on tabId: 1899282415</li>\n<li>Available tabs:\n• tabId 1899282384: &quot;Sloppy Models &amp; Potential Transfer - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d3252c-a902-84ec-8000-09e439b2dbef?chat_enter_method=home\">https://www.kimi.com/chat/19d3252c-a902-84ec-8000-09e439b2dbef?chat_enter_method=home</a>)\n• tabId 1899282383: &quot;GLIM Briefing Follow-Up - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d32582-9fd2-8d82-8000-09e4e73fcf9b?chat_enter_method=change_model\">https://www.kimi.com/chat/19d32582-9fd2-8d82-8000-09e4e73fcf9b?chat_enter_method=change_model</a>)\n• tabId 1899282382: &quot;Information-Theoretic Model Compression - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d3288e-0aa2-8d6b-8000-09e4a6fbe0d1?chat_enter_method=home\">https://www.kimi.com/chat/19d3288e-0aa2-8d6b-8000-09e4a6fbe0d1?chat_enter_method=home</a>)\n• tabId 1899282409: &quot;Kimi Deep Research | From Investigation to Insights&quot; (<a href=\"https://www.kimi.com/deep-research\">https://www.kimi.com/deep-research</a>)\n• tabId 1899282394: &quot;Materials Benchmarking Simpson&#39;s Paradox - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d32652-ec92-8250-8000-09e4bb02cdcc?chat_enter_method=home\">https://www.kimi.com/chat/19d32652-ec92-8250-8000-09e4bb02cdcc?chat_enter_method=home</a>)\n• tabId 1899282397: &quot;Materials Informatics Funding Landscape - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d32876-40b2-8606-8000-09e4e29c29fd?chat_enter_method=home\">https://www.kimi.com/chat/19d32876-40b2-8606-8000-09e4e29c29fd?chat_enter_method=home</a>)\n• tabId 1899282406: &quot;RG Coarse-Graining in MD - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d32834-ded2-8ea6-8000-09e4c33a8a29?chat_enter_method=home\">https://www.kimi.com/chat/19d32834-ded2-8ea6-8000-09e4c33a8a29?chat_enter_method=home</a>)\n• tabId 1899282412: &quot;GLIM molecular dynamics meta-analysis platform - Claude&quot; (<a href=\"https://claude.ai/chat/3a5c0904-db84-4c96-b423-a7def2513b18\">https://claude.ai/chat/3a5c0904-db84-4c96-b423-a7def2513b18</a>)\n• tabId 1899282415: &quot;TDA in Interatomic Error Landscapes - Kimi&quot; (<a href=\"https://www.kimi.com/chat/19d32e07-de52-8c63-8000-09e4189ac824\">https://www.kimi.com/chat/19d32e07-de52-8c63-8000-09e4189ac824</a>)</li>\n</ul>\n"}