Completed Public Health & Healthcare Pregnancy, Children & Inherited Conditions

Data mining epidemiological relationships: integration of causal analysis with published evidence

In plain English

AI plain-English summary

A single graph database will map hundreds of known and suspected links between risk factors and diseases, allowing researchers to search for hidden causal connections across the entire network. Epidemiologists typically test one risk factor against one disease at a time. This narrow focus misses the bigger picture: that smoking, diet, exercise, and genetics interact in complex webs, and that an intervention targeting one risk factor may have unintended side effects on others. The researchers will build a purpose-built “graph” database that integrates published causal relationships with biological data—molecular pathways, drug targets, and disease outcomes. They will then develop computational methods to mine this network for novel causal risk factors and potential interventions. If successful, the open-access software platform will let other researchers search the integrated datasets for their own questions, accelerating discovery without requiring each lab to rebuild the same connections. This is primarily a tool-building and data-integration project. It does not directly change clinical practice or public health guidance today, but it could help identify which risk factors matter most for a given disease, and flag unintended consequences of proposed interventions—making future epidemiological studies more efficient and their conclusions more robust.

View original technical description
Causal inference in epidemiology focuses on identifying the risk factors that cause disease. Established approaches focus on specific risk factors that may impact on specific diseases. However, the wealth of biomedical data that now exist enable us to assess the causal relationships between a broad network of risk factors and diseases. By considering a much wider network of such relationships we will establish the relative importance of different risk factors and the potential side-effects of interventions that target those risk factors. We will also integrate biological data (eg molecular pathways, drug targets) with causal relationships to enable us to understand the molecular mechanisms that lead to disease, and identify potential pharmaceutical and public health interventions. These data and relationships will be combined in a purpose-built “graph” database and methods will be developed to mine for novel causal risk factors and potential interventions. The data that we collate for our research within this programme will have wide-reaching value to the research community. We will provide an open and accessible software platform for other researchers to search and use the various datasets we have integrated for their own research.

View the original record at the funder ↗

Researchers

Tom Gaunt (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Data mining epidemiological relationships
Investigating methods for inferring causality from observational data: an application to longitudinal cohort data
Strategy for analysing epidemiological data involving genetic, endogenous, environmental factors and their interactions
Investigating comorbidities with causal networks
Causal Inference from Partial Statistical Information

Original classification

Intramural

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.