Active Public Health & Healthcare Mathematics & Statistics

Developing guidance for multiple imputation

In plain English

AI plain-English summary

Missing data silently corrupts the results of medical trials and social surveys, leading to policies and treatments based on flawed evidence. When people drop out of a study or skip questions, the remaining data can paint a misleading picture unless analysed correctly. Multiple imputation (MI) is a statistical technique that fills in those gaps by predicting missing information from what is known about participants, but current guidance on how to use it is often outdated, contradictory, or simply absent for common scenarios like comparing patient survival under different treatments. This project will produce clear, up-to-date guidance on when and how to apply MI, co-developed with journal editors and policy advisors who assess statistical methods. The team will create free software tools and worked examples that make best practice easy to follow, and will run workshops to train researchers. By standardising how missing data is handled and requiring analysts to document their decisions and code, the guidance will make studies more reproducible and transparent. If widely adopted, it could improve the reliability of evidence that underpins healthcare guidelines, public health interventions, and social policy in the UK and globally.

View original technical description
Missing data is a problem across all studies in health and social sciences, randomised trials, and social surveys. Information may be missing for a variety of reasons – e.g. because people don't want to answer some questions, or forget to give some information, or drop out completely – and often the reasons for the missing data are not known. The way the available information is analysed must be chosen carefully, or the results of the study will be wrong (“biased”), or less precise than they should be, or both. This could mean that policies, therapies, or interventions are based on incorrect evidence. Multiple imputation (MI) is an analysis strategy that can correct the bias due to missing data. In MI, the information we do know about people in the study (e.g. details of their previous health, their age, and so on) is used to predict ("impute") the missing information. Whether this technique is successful depends on why the information is missing in the first place, and how well it can be predicted. There are some guidelines for carrying out MI in specific circumstances, but these are often outdated, incorrect, or complex and hard to follow. There are also commonly-encountered scenarios – e.g. comparing patient survival under different treatment strategies – where guidelines for MI do not exist. Different studies use MI in different ways, and do not usually document what was done - so it is hard to replicate analyses, or to see if analysts have followed best practice. We aim to develop guidance on when and how to carry out MI. We will focus on aspects of MI for which guidance is particularly lacking. This will be useful for anyone who uses incomplete data, but will be especially aimed at those who may have relatively little formal training in statistical analysis of missing data. We will co-produce our guidance with experts in statistical methods, including those who assess statistical research submitted to medical journals, and those who already provide policies and advice for users of such data. This will ensure our guidance is up-to-date, relevant, accessible, and addresses areas of need. We will provide our guidance in many formats. This will include freely-available software tools that make it easy to apply each part of our guidance, and examples that show exactly how our tools can be used. Our guidance will ensure users are following best practice and - by providing documented decisions and code - it will increase reproducibility and transparency of analyses. We will run focus groups with researchers to help us develop and refine our guidance. We will include the guidance, and the tools to implement it, in workshops and courses on how to deal with missing data, in order to reach as many people as possible. Our guidance will be useful for all types of study – including cohort studies, randomised trials, and surveys. We will use our links with other researchers, those working in health and social settings, and non-academic agencies to ensure that our guidance is widely used. Thus it will have the potential to improve the level of evidence informing policy and practice in health, medicine, and beyond, in the UK and worldwide.

View the original record at the funder ↗

Researchers

Elinor Curnow (Principal Investigator)James Carpenter (Co-Investigator)Jon Heron (Co-Investigator)Kate Tilling (Co-Investigator)Rosie Cornish (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Development of miDOC: an expert system and methodology for multiple imputation
Multiple imputation by chained equations for data that are missing not at random: methods development for randomised trials and observational studies
Developing and disseminating robust methods for handling missing data in epidemiological studies
Debiased machine learning for missing data
Artificial Intelligence for Missing Data Imputation in Electronic Medical Records

Original classification

Research and Innovation

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.