Doctors are prescribing diabetes drugs to patients who were not included in the original clinical trials, creating uncertainty about whether those drugs still work as intended. This PhD project will test whether machine learning can extract reliable answers from routine GP records—data that is messy and prone to hidden biases—to fill that gap. The core problem is that the landmark trials for SGLT2 inhibitors, which protect the heart in people with type 2 diabetes, excluded many patients now receiving the drugs. Trials also cannot detect whether effects differ by age, sex, or other health conditions. The researcher will mimic a key trial using UK primary care data, then compare traditional statistical methods with machine learning techniques—including natural language processing of doctors’ notes—to see which best removes confounding. A final step will use causal forests and meta-learners to identify which patient subgroups benefit most or least. If successful, this work could give clinicians and regulators real-world evidence on drug effects for the patients they actually see, not just the ones in trials. It may also establish best practices for analysing observational data in pharmacoepidemiology, making future real-world studies more trustworthy.
View original technical description
Research question: How can machine learning (ML) be used to develop our understanding of real-world effects of anti-diabetes drugs? Background: Cardiovascular diseases (CVD) are a major cause of morbidity and mortality among people with diabetes. To address this large burden of disease, sodium-glucose cotransporter-2 inhibitors (SGLT2i) have been introduced as a first-line therapy for people with type 2 diabetes at risk of atherosclerotic CVD. This is based on the findings of large randomised controlled trials (RCTs). However, these therapeutics are now being prescribed for people who would not have been included in the original RCTs and for wider therapeutic indications. Similarly, RCTs are under-powered to detect meaningful subgroup effects. As a result, there is uncertainty as to how the findings of RCTs translate into real-world treatment effects. There is growing interest in leveraging routinely collected observational data to better understand real-world drug effects, but the use of such data is limited by confounding. Several methods, including machine learning (ML), have shown promise at reducing confounding in observational studies, but the practical implementation of such methods to real-world data is limited. This PhD aims to translate developments in causal inference and ML into clinical insights relating to anti-diabetic agents, with three aims: Aims: 1. Compare ML and traditional methods for covariate adjustment in a target trial framework. 2. Determine the real-world effects of anti-diabetic agents in a target trial using ML methods. 3. Apply ML and causal inference methods to study real-world heterogeneous treatment effects and consider how such insights might inform patients and clinicians. Methods: The foundation of this PhD will be a target trial of the EMPA-REG outcome RCT, which was a seminal study that helped to establish the cardioprotective effects of SGLT2i in patients with type 2 diabetes and established atherosclerotic CVD. This study will be emulated in an observational setting, using UK primary care data from The Health Improving Network. Employing this target trial, we will compare traditional and advanced covariate adjustment methods to determine the optimal strategy for real-world analysis. Firstly, we will explore if including symptom covariates, derived from a natural language processing algorithms applied to free text notes, reduces residual confounding. Secondly, we will estimate and utilise high-dimensionality propensity scores and ML-derived propensity scores, implemented using targeted maximum likelihood estimation and double machine learning. Finally, we will explore heterogeneous treatment effects according to key demographics and comorbidities using ML methods such as meta-learners and causal forests. The real-world estimates will be compared to the original RCT data. Patient and public panels will be conducted to inform real-world data analysis and ensure the output of the PhD addresses the priorities of people living with diabetes. Impact: This work will enrich the evidence-base for anti-diabetes drugs with insights from real-world data - informing patients, clinicians and regulatory decisions. It may motivate the conduct of confirmatory clinical trials. This PhD will also inform best practices for real-world data analysis by studying the utility of ML methods in pharmacoepidemiology research and promoting the robust analysis of observational data.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know