Active Genetics & Molecular Biology Cells, Biochemistry & Physiology

Machine Learning Methods for Dynamic Gene Regulation and Personalised Epigenomics

In plain English

AI plain-English summary

A patient’s blood sample could one day reveal what is happening in their brain or liver, without needing a biopsy. The researcher is building machine learning tools to read the epigenome—the chemical switches on DNA that turn genes on or off depending on cell type, environment, and disease state. Current methods struggle because the epigenome differs across tissues and is noisy to measure. The team’s tool, eDICE, uses existing data to predict epigenomic profiles for hard-to-access tissues from blood samples. A second tool, DecoDen, separates true biological signals from measurement noise. Together, they aim to make personalised epigenomics practical. If successful, this could transform how doctors track slow-developing diseases like cancer or neurodegeneration, by detecting early epigenetic changes without invasive procedures. The second part of the project is fundamental science: the team will experimentally degrade epigenetic remodelers in cells and use a new causal discovery framework to map how these disruptions ripple across the genome. This deeper understanding of how cells maintain—or lose—their identity could eventually inform drug targets for cancers driven by mutations in epigenetic regulators.

View original technical description
The long-term goal of my work is to innovate methodologies for understanding epigenetic processes that shape developmental dynamics and patient-specific disease manifestation. To achieve this, I combine experimental design, data generation, and machine learning. In the early 2000s, the Human Genome Project (HGP) promised to transform disease understanding by mapping the human genome. Large-scale efforts like genome-wide association studies (GWAS) revealed statistical links between genetic variants and disease. However, GWAS struggles to identify causal factors, as most common diseases are polygenic and multifactorial, influenced by genes, environment, and lifestyle. Epigenetics bridges the gap between static genetic information and dynamic biological contexts, offering crucial insights into disease mechanisms. Unlike the genome, the epigenome is cell-type specific, with significant variation in histone post-translational modifications (PTMs) across tissues. This diversity creates unmet demand for experimental data. To address this, we developed eDICE, a computational tool leveraging existing reference datasets to impute unmeasured epigenomic profiles. Using transfer learning, eDICE predicts individual-specific epigenomic landscapes, enabling tissue-specific predictions for unobserved tissues. While eDICE achieved high accuracy in an international benchmark, independent validation highlighted variability in epigenomic measurements due to intrinsic noise. To improve reliability, we developed DecoDen, a method integrating replicates, multi-ChIP, and control sequencing data to separate true signals from noise. Combining DecoDen and eDICE will enhance the specificity of individual- and tissue-specific pattern prediction, enabling translational epigenomics, such as inferring hard-to-access tissue profiles (e.g., brain or liver) from blood samples. Analyzing individual-specific epigenomic patterns is vital, as these patterns act as cellular memories of past transcriptional events and may encode essential information about slowly developing disease pathways. At the same time, epigenomic mechanisms provide robustness against unintended reprogramming, maintaining stability. This dual role is evident in the prevalence of mutations in epigenetic modulators, such as ARID1A, which are frequently altered in diverse cancers. Paradoxically, these modulators safeguard genomic integrity while also enabling cellular plasticity during development. In the second part of this project, we will investigate immediate cellular responses to experimentally induced degradation of epigenetic remodelers in cell cultures, uncovering their specific mechanisms of action. Traditional gene regulatory networks have provided insights into gene perturbations, knockouts, and overexpression but struggle with epigenomic interventions, which disrupt chromatin structure genome-wide. To address this, we will introduce Causal Discovery from Conditionally Stationary Time Series to model epigenomic perturbations. While causal discovery is challenging for AI systems, especially with genome-wide disruptions, adapting these methods for conditionally stationary data will create a robust framework for understanding dynamic causal relationships driving epigenomic changes and their effects.

View the original record at the funder ↗

Researchers

Gabriele Schweikert (Principal Investigator)Tom Owen-Hughes (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Machine Learning Methods to Re-annotate Histone Modifications with Locus-specific Functional Classification
Novel Machine Learning Techniques to Elucidate Function and Dynamics of Epigenomic Mechanisms
Machine-learning to create predictive models of genetic regulation
Development of software to model multi-modal genomic data as an integrated system: application to understanding the gene regulatory landscape
Development of a multilevel and mixture-model framework for modelling epigenetic changes over time (resubmission)

Original classification

Fellowship

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.