Recipient organisationCardiff UniversitySource-published name: Cardiff University
Funding£521K
PeriodMay 2025 — May 2028
In plain English
AI plain-English summary
Statisticians are building a new mathematical toolkit that lets researchers analyse medical data on its original scale—without first transforming it into something artificial. Most statistical models force data into unnatural shapes, like converting survival times into logarithms, which makes results harder for doctors to interpret. Regression by composition keeps everything on the original scale of measurement—years, blood pressure readings, tumour sizes—so the numbers clinicians see are the numbers they can act on. The team will develop the underlying theory, release open-source software, and test the approach in clinical trials, epidemiology, and Mendelian randomisation studies. If successful, this could change how medical statisticians teach their subject and how trial results are communicated to doctors. Because the method is flexible enough to handle high-dimensional machine learning models while remaining interpretable, it could also serve as a bridge between black-box AI predictions and the transparent reasoning regulators and clinicians demand. The project is primarily methodological—it builds fundamental statistical theory—but its direct applicability to real-world data means practical impact could follow quickly.
View original technical description
Our research group has recently proposed regression by composition, a flexible toolkit for building and understanding statistical models. Regression by composition always takes place on the original scale of the data, so offers enormous potential for clinical insight and better decision-making. We seek to enrich the mathematical theory of regression by composition, to release open-source software that expands its applicability, and to evaluate its effectiveness in biomedical science. We will... ...engage specialists in clinical trials, epidemiology, biostatistics and statistical computing to ensure regression by composition is widely applicable in biomedical research, acceptable to its scientists, and accessible to its analysts. More broadly, regression by composition affords an opportunity for reimagining teaching and training in medical statistics. ...assemble a computational engine for fitting regressions by composition. A major objective is to provide general-purpose routines suitable for specialised or user-supplied model components, while taking advantage of exact mathematical results where these are known. ...design flexible longitudinal regressions by composition that are compatible with structural nested models, and congenial for instrumental variable regression or Mendelian randomisation. Because interventions can influence health trajectories in surprising and subtle ways, we will guide scientists towards transportable findings, combining empirical evidence with careful application of mechanistic, first-principles reasoning. ...employ assumption-lean regressions by composition as tools for explainable artificial intelligence. Regression by composition is general enough to express high-dimensional machine learning models; with the addition of tuning parameters controlling regularisation, we will give users the ability to select from a spectrum of accurate and easy-to-understand models for classification, prediction or control.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know