Active Public Health & Healthcare Genetics & Molecular Biology

Data mining epidemiological relationships

In plain English

AI plain-English summary

Everyday lifestyle choices—like diet, exercise, or smoking—can raise or lower a person's risk of common diseases, but pinning down which factors actually cause harm, rather than just being correlated with it, has been notoriously difficult. This research programme is building automated data mining tools that use genetic information to separate true causal risk factors from misleading associations. The team is also constructing a "knowledge graph"—a vast, interconnected map of biomedical evidence—that links genetic data, disease mechanisms, and drug effects. This allows them to systematically search for new drug targets, predict side effects, and identify existing drugs that could be repurposed for other conditions. If successful, the tools and knowledge graph will be made openly available to the global research community. This could accelerate the discovery of modifiable risk factors for diseases like heart disease or diabetes, and speed up the identification of new treatments—without requiring new clinical trials for every candidate. The work is primarily methodological and fundamental in nature, but similar open-source analytical tools have previously transformed how researchers use large-scale population data to improve public health.

View original technical description
We aim to develop and use cutting edge data mining tools to identify risk factors that cause common diseases and potential drug targets that could prevent or treat these diseases. Methods developed within the MRC Integrative Epidemiology Unit use genetic data to help identify lifestyle risk factors that could be modified to reduce the risk or impact of disease, and can also identify potential drug targets. This programme is developing tools and databases to automate this type of analysis and apply it to large-scale population datasets to help us discover new ways to prevent and treat disease. We are also combining the evidence from these analyses with other types of biomedical information in a “knowledge graph” to enable us to investigate the mechanisms underlying disease, identify new targets for treatment or prevention, predict side effects of drugs and identify opportunities to repurpose existing drugs for other diseases. The methods, software and knowledge graph we are developing are made openly available to the research community to maximise their potential to improve population health.

View the original record at the funder ↗

Researchers

Tom Gaunt (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Data mining epidemiological relationships: integration of causal analysis with published evidence
Understanding disease through environment-wide association studies
Mendelian Randomisation
Bayesian Discovery of Regression Structures: a tool kit for genetic epidemiology and integrative genomics analyses
Managing and exploiting high dimensionality in genetic epidemiology

Original classification

Intramural

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.