Completed Mathematics & Statistics Computing & AI

Sample complexity for Gaussian graphical models and liftability of geometric structures

In plain English

AI plain-English summary

Fitting a high-dimensional statistical model to a small dataset is like trying to triangulate a location with too few satellites—the maths often fails. This project tackles that exact problem, using the geometry of scaffold-like structures made of bars and joints to determine when a statistical model can be reliably estimated from limited data. In fields such as neuroscience, structural biology, and archaeology, collecting many data points is expensive or impossible, yet researchers still need to draw robust conclusions from what little they have. The core question is: which patterns of dependencies among variables allow a maximum-likelihood estimate to exist when the sample size is smaller than the number of variables? The answer, the team has discovered, lies in an unexpected link to rigidity theory—the mathematics of whether a framework of bars can be “lifted” into a higher dimension without changing bar lengths. This is fundamental science. If successful, it will provide a geometric toolkit for deciding, in advance, whether a given dataset can support a particular statistical model. That could eventually improve the reliability of models used in brain imaging, ancient DNA analysis, or any field where data is scarce but conclusions carry weight.

View original technical description
A central task in data science is to fit a large multivariate normal model from fewer datapoints than variables. The reason is that, in many applications such as neuroscience, structural biology or archaeology, data is either expensive or even impossible to obtain. A widely-used approach is to assume sparsity from some known conditional independences among pairs of variables and then use maximum likelihood estimation. This project will establish new geometric and combinatorial approaches to the question of which independences allow for maximum-likelihood estimation from a fixed sized sample. More generally, the project will study which linear constraints on the inverse covariance imply the existence of a maximum likelihood estimator for a fixed size sample—a question that involves a subtle interplay between geometry and structural graph theory. The project will exploit recently-discovered links between such data complexity problems and rigidity theory, a field of discrete geometry at the intersection of combinatorics, algebra, and geometry. Consider a scaffold-like structure composed of stiff bars linked at universal joints, with the connectivity of the bars modelled by a graph. Spectral properties of a weighted Laplacian of the structure’s graph control whether the structure can be “lifted” to another structure with maximal affine span and the same bar lengths. Whether a typical structure in a given dimension can be lifted determines whether the maximum likelihood estimator for an associated Gaussian graphical model exists, providing a geometric avatar for the original statistical problem. A key feature of the project is to initiate an in-depth and sustained exploration of the interplay between linear concentration models and the geometry of rigid structures. The main goals are to: (1) characterise the linear concentration models that have a maximum likelihood estimator for each number of samples; and (2) develop geometric and combinatorial rigidity of linear measurement processes on squared edge lengths and characterise the associated algebraic matroids.

View the original record at the funder ↗

Researchers

Anthony Nixon (Co-Investigator)Louis Theran (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Geometric deep learning for likelihood-free statistical inference
FORGING: Fortuitous Geometries and Compressive Learning
Exact scalable inference for coalescent processes
Statistical physics insights on inference and learning of structured and heavy-tailed datasets
Algorithms, Dynamics and Connections with Phase Transitions

Original classification

Research and Innovation

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.