Completed Mathematics & Statistics Computing & AI

Statistical inference for high-dimensional heterogenous data

In plain English

AI plain-English summary

Every month, economic data can suddenly shift because a government changes a policy or a trade deal collapses—and statisticians need to spot exactly when that happened, even when the dataset contains thousands of variables. This project tackles a fundamental problem in modern data analysis: high-dimensional datasets—where the number of measured variables rivals the number of observations—often contain hidden change points, moments when the underlying pattern generating the data abruptly alters. Existing methods struggle to detect these change points reliably, especially when the number of shifts is unknown. The researcher will develop a technique called weighted empirical risk minimisation, which builds on standard statistical tools like LASSO and logistic regression but adds Bayesian weights that encode prior knowledge—for instance, that change points must be at least a certain distance apart. The method is computationally efficient, as easy to implement as the underlying estimator, and comes with rigorous mathematical guarantees on its accuracy. If successful, the work will give researchers in medicine, economics, and biology a practical, ready-to-use toolkit for analysing heterogeneous data. It is primarily fundamental science, advancing statistical theory and forging new connections between high-dimensional statistics, information theory, and signal processing. Similar foundational work on sparse regression and convex optimisation has already transformed fields from genomics to machine learning.

View original technical description
Heterogeneity is a common feature of large, high-dimensional datasets. A simple form of heterogeneity in data ordered by time is a change in the data generating mechanism at certain unknown time instants. For example, monthly economic data may be affected by changes in government policy or in import/export regulations of other countries. A key question is: how can we fit accurate statistical models for high-dimensional data in the presence of such change points? Change points are the time instants where the underlying generative mechanism changes. The challenge in fitting regression models is especially acute with modern datasets, where even the number of change points is often unknown. The predictions of a model may be highly misleading if the change points are not estimated accurately. This project addresses the challenge of detecting change points in high-dimensional regression models, where the data dimension is large and comparable to the sample size. We will develop a novel technique for change point regression based on weighted empirical risk minimization (weighted ERM). Our approach allows prior information on the changepoints to be encoded via Bayesian weights. For example, the weights could encode a minimum separation between the change points. The weighted ERM approach allows us to adapt standard convex penalized estimators (such as LASSO and sparse logistic regression) for change point detection. The weighted ERM estimator is as easy to implement as the underlying convex estimator, with complexity of the same order. We will demonstrate the performance of the technique using datasets drawn from fields ranging from the medical sciences to econometrics. Moreover, we will establish rigorous guarantees on the performance of weighted ERM on high-dimensional data, including a computable posterior distribution on the change points. The weighted ERM technique developed in this project is practical, and can be applied for change point detection in a broad class of generalized linear models, including the widely used logistic, probit, and Poisson models. Generalized linear models are a workhorse of statistics, and are widely used for regression and classification in biology, medicine, economics, and many other fields. The project will provide a powerful, easy-to-use, set of tools for analyzing the heterogenous, high-dimensional datasets often encountered in these areas. Drawing on ideas and techniques from high-dimensional statistics, information theory, and signal processing, the project will establish novel theoretical and algorithmic connections between these areas. We expect that it will open the way to a much broader investigation of heterogeneity in high-dimensional data.

View the original record at the funder ↗

Researchers

Ramji Venkataramanan (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Change-point detection for high-dimensional time series with nonstationarities
Change-point analysis in high dimensions
Bayesian High-Dimensional Time series models with applications to Macroeconomic and Financial data
DMS-EPSRC: Change Point Detection and Localization in High-Dimensions: Theory and Methods
Statistical inference for high-dimensional data

Original classification

Research and Innovation

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.