Statisticians are building a unified toolkit to solve problems where the standard mathematical formula for drawing conclusions from data simply cannot be computed. This matters because many of today’s biggest scientific and commercial questions—tracking disease outbreaks, analysing massive genetic datasets, modelling ecosystems, or understanding patterns in online commerce—produce data so large or complex that existing likelihood-based methods grind to a halt. The 1990s revolution in computational statistics opened up likelihood inference to many fields, but high-dimensional problems and enormous datasets now routinely exceed those methods’ limits. If the Ilike project succeeds, it will deliver a coherent, practical framework that combines recent breakthroughs—such as particle Markov chain Monte Carlo, approximate Bayesian computation, and adaptive Monte Carlo—with modern multi-core computing hardware. This would let researchers in genetics, epidemiology, ecology, and bibliometrics draw reliable statistical conclusions from problems that are currently intractable. The work is primarily fundamental methodology, but its direct application to real-world data from the outset means the impact could be felt wherever complex, large-scale inference is needed.
View original technical description
In most statistical contexts, it is recognised that inference methodology based on the likelihood function are usually methods of choice. However such methods are not always easy to implement. For instance, in complex problems often with massive data sets, it can sometimes be completely impossible to even evaluate the likelihood function. The computational statistics revolution of the 1990s provided powerful methodology for carrying out likelihood-based inference, including Markov chain Monte Carlo methods, the EM algorithm, many associated optimisation techniques for likelihoods, and Sequential Monte Carlo methods. Although these methods have been and are highly successful in making likelihood-based inference accessible to a wide range of problems from virtually every area of science and technology, we now have a far better understanding of their limitations, for example in high-dimensional problems and for massive data sets. Thus many challenging statistical inference problems of the 21st century cannot be addressed using existing likelihood-based methods. Examples which motivate the current project come from genetics, genomics, infectious disease epidemiology, ecology, commerce, and bibliometrics. However there have been various recent breakthroughs in computational and statistical approaches to intractable likelihood problems, including pseudo-marginal and particle MCMC, likelihood-free methods such as Approximate Bayesian Computation, composite and pseudo-likelihoods, new simulation methods for hitherto intractable stochastic models, and adaptive Monte Carlo methods. These advances coupled with developments in multi-core computational technologies such as GPUs, have enormous potential for extending likelihood methods to meet the most difficult challenges of modern scientific questions. All these new areas have demonstrated considerable promise. However, a unified approach is required to deliver a step change in our capability to implement likelihood. The Ilike workplan will initially involve the investigation of 10 projects all combining two or more of the about breakthrough areas. The second half of the project will involve a similar workplan informed by the outcomes of the initial projects.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know