Completed Mathematics & Statistics Computing & AI

StatScale: Statistical Scalability for Streaming Data

In plain English

AI plain-English summary

London's Oyster card system and the Square Kilometre Array telescope both generate torrents of data that arrive continuously over time, and standard statistical methods struggle to keep up. The problem is that existing statistical algorithms, designed for smaller datasets, fail in two ways when faced with massive data streams. First, the mathematical models at their core are approximations, and this "model error" becomes dangerously large when processing billions of data points. Second, the fastest computational shortcuts often sacrifice statistical accuracy, forcing a trade-off that has no good solution. This programme grant aims to build a new generation of statistical methods that handle both issues simultaneously, focusing on detecting when the structure of a data stream changes—a problem identified as one of seven key challenges for Big Data. If successful, the work could improve real-time systems that quietly underpin modern life: transport networks adjusting to passenger flows, energy grids balancing supply and demand, or financial systems spotting anomalies in transaction streams. The research is fundamentally about the mathematics of inference at scale, but its practical payoff would be more reliable decisions made from the flood of data that sensors, smartphones, and scientific instruments already produce.

View original technical description
We live in the age of data. Technology is transforming our ability to collect and store data on unprecedented scales. From the use of Oyster card data to improve London's transport network, to the Square Kilometre Array astrophysics project that has the potential to transform our understanding of the universe, Big Data can inform and enrich many aspects of our lives. Due to the widespread use of sensor-based systems in everyday life, with even smartphones having sensors that can monitor location and activity level, much of the explosion of data is in the form of data streams: data from one or more related sources that arrive over time. It has even been estimates that there will be over 30 billion devices collecting data streams by 2020. The important role of Statistics within "Big Data" and data streams has been clear for some time. However the current tendency has been to focus purely on algorithmic scalability, such as how to develop versions of existing statistical algorithms that scale better with the amount of data. Such an approach, however, ignores the fact that fundamentally new issues often arise when dealing with data sets of this magnitude, and highly innovative solutions are required. Model error is one such issue. Many statistical approaches are based on the use of mathematical models for data. These models are only approximations of the real data-generating mechanisms. In traditional applications, this model error is usually small compared with the inherent sampling variability of the data, and can be overlooked. However, there is an increasing realisation that model error can dominate in Big Data applications. Understanding the impact of model error, and developing robust methods that have excellent statistical properties even in the presence of model error, are major challenges. A second issue is that many current statistical approaches are not computationally feasible for Big Data. In practice we will often need to use less efficient statistical methods that are computationally faster, or require less computer memory. This introduces a statistical-computational trade-off that is unique to Big Data, leading to many open theoretical questions, and important practical problems. The strategic vision for this programme grant is to investigate and develop an integrated approach to tackling these and other fundamental statistical challenges. In order to do this we will focus in particular on analysing data streams. An important issue with this type of data is detecting changes in the structure of the data over time. This will be an early area of focus for the programme, as it has been identified as one of seven key problem areas for Big Data. Moreover it is an area in which our research will lead to practically important breakthroughs. Our philosophy is to tackle methodological, theoretical and computational aspects of these statistical problems together, an approach that is only possible through the programme grant scheme. Such a broad perspective is essential to achieve the substantive fundamental advances in statistics envisaged, and to ensure our new methods are sufficiently robust and efficient to be widely adopted by academics, industry and society more generally.

View the original record at the funder ↗

Researchers

Idris Eckley (Principal Investigator)John Aston (Co-Investigator)Paul Fearnhead (Co-Investigator)Rajen Shah (Co-Investigator)Richard Samworth (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Statistical Foundations for Detecting Anomalous Structure in Stream Settings (DASS)
Dynamic Scaling of Distributed Dataflows under Uncertainty
Statistical methodology and theory for the Big Data era (Ext.)
Novel statistical methods for detecting anomalies in data streams.
Emerging Geometries for Statistical Science: Articulating the Vision

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.