Active Mathematics & Statistics Public Health & Healthcare

Statistical methodology for the analysis of Electronic Health Records

In plain English

AI plain-English summary

Electronic health records contain millions of patient histories, but the data is riddled with gaps and errors that can skew results if analysed with standard statistical tools. This matters because researchers and doctors increasingly rely on these records to answer basic questions—how common a disease is, who gets it, and how it progresses. The data was never designed for research; it was collected for clinical care. Simple methods that ignore missing or incorrect entries can produce misleading conclusions, potentially leading to wrong screening recommendations or treatment decisions. The researcher is developing new statistical methods that handle missing data, combine information from different sources, and account for uncertainty. They are also building software tools so other researchers can apply these methods easily. If successful, this work will improve the accuracy of any study that uses electronic health records. That includes helping doctors decide when to invite people for screening, guiding treatment choices, and giving a clearer picture of how diseases unfold over time. The impact is not dramatic or visible—it is a quiet improvement to the analytical backbone of modern epidemiology and clinical decision-making.

View original technical description
We often want to know how common a disease is (prevalence), how often people get a disease (incidence) and who is most likely to get a disease. We have previous patients’ electronic health records on computers which we can use to answer these questions. This form of data has many benefits as it is quick and cheap to collect whilst covering diverse populations. However, there can be problems as the data was not originally intended for research and contains missing and incorrect data. If we use simple statistical methods that ignore these issues, we can end up with misleading results. I propose to develop new statistical methods that can handle the missing data, combine different data sources, and can take into account inaccurate or uncertain data. I will develop tools which will encourage and enable other researchers to use the methods as well. I expect my work to improve the accuracy of research which uses electronic medical records. In particular, the methods will help doctors know when to invite people for disease screening, help guide doctors and patients on treatment decisions and provide a greater insight on disease progression.

View the original record at the funder ↗

Researchers

Matilda Pitt (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Improving decisions on what to focus on in research using large datasets
Developing and disseminating robust methods for handling missing data in epidemiological studies
Providing more accurate predictions of colorectal cancer prognosis: development and application of novel methodology when using electronic health record data for disease prognosis
Statistical methods for developing, assessing and validating risk prediction models in multiple epidemiological studies
Handling missing data in large electronic healthcare record datasets

Original classification

PhD Studentship (Basic)

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.