Active Public Health & Healthcare Mathematics & Statistics

RESSOLVE-HD: Regularized Estimation for Stable Solutions to Overcome Selection Bias in Health Data

In plain English

AI plain-English summary

When people skip HIV tests, the resulting gaps in data can make infection rates look lower than they really are. This project builds statistical tools to correct for that kind of hidden bias. The problem is that health surveys often suffer from missing data—people opt out, refuse tests, or drop out of studies. If those missing people differ systematically from those who stay, the final numbers are skewed. This is especially acute in HIV prevalence studies, where low participation can distort national estimates and mislead policymakers. Current statistical methods exist to correct for this, but they rely on identifying "exclusion restriction variables"—factors that affect participation but not the outcome itself—which is notoriously difficult. The researchers aim to develop stable, reliable techniques that do not require perfect knowledge of those variables. They will determine the sample sizes and effect sizes needed for robust corrections, handle missing data in the covariates themselves, and build accessible software so that non-specialists can use the methods. If successful, the work will give epidemiologists and public health officials more accurate estimates of disease burden, particularly for HIV, and could be applied to other health surveys where non-ignorable missing data threatens the validity of conclusions.

View original technical description
In health research, missing data can skew results, especially when data are missing for reasons related to unobserved values (non-ignorable missing data). This bias can occur due to sample selection, where individuals opt out of participation, affecting the representativeness of the sample. For example, in HIV (Human Immunodeficiency Virus) prevalence studies, low participation in testing can distort our understanding of HIV rates in the population. Sample selection models (SSMs) help address this but identifying exclusion restriction variables (ERVs) – factors affecting study participation but not outcomes – is challenging. The aim of this project is to develop novel statistical techniques for overcoming this challenge. We will focus on establishing the necessary sample sizes and effect sizes for robust SSM applications, ensuring reliability for medical researchers and decision- makers. Our approach will involve developing stable variable selection methods for SSMs, handling missing data in covariates, and creating accessible software for wider adoption. Using both simulated and real-world datasets, including those from the Demographic and Health Surveys and nationally representative studies, we aim to provide practical solutions for addressing missing data in health research.

View the original record at the funder ↗

Researchers

Emmanuel Ogundimu (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Semiparametric Sample Selection Models with Applications in Biostatistics, Economics and Environmetrics
Robust statistical methods and statistical diagnostic techniques for multivariate longitudinal and survival data in health research
Developing and disseminating robust methods for handling missing data in epidemiological studies
Investigating transportability of cancer detection models across datasets and time using population-wide electronic health data
Statistical Learning and Adaptive Observation in Clinical Prediction: Methodology and Applications

Original classification

Wellcome Accelerator Awards

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.