Completed Genetics & Molecular Biology Cells, Biochemistry & Physiology

Machine Learning Methods to Re-annotate Histone Modifications with Locus-specific Functional Classification

In plain English

AI plain-English summary

Every cell in the human body contains the same DNA, yet a nerve cell and a blood cell look and behave completely differently—epigenetic marks are the chemical switches that tell each cell which genes to use and which to ignore. This project tackles a fundamental gap in biology: scientists can measure these epigenetic marks, but they do not fully understand what they mean. The "histone code"—the idea that chemical modifications on histone proteins control gene activity—has remained a mystery for decades. The researcher will build machine learning tools to uncover how these marks are written, read, and erased in living cells, then test those predictions with lab experiments that rapidly degrade specific epigenetic enzymes. If successful, this work will provide a functional map of the epigenome—showing not just where marks sit, but what they actually do. This is fundamental science, not an applied technology. However, because faulty epigenetic regulation is linked to cancers and developmental disorders, a clearer understanding of how these switches work could eventually guide new strategies for diagnosing or treating diseases where cells lose their identity.

View original technical description
The human body contains about 200 different cell types, e.g. nerve or blood cells, each with their specific appearances and functions. To carry out their proper roles, they all execute different sets of genetic programs while containing an identical copy of the complete genomic instructions (the DNA), which is passed down from a single parent cell. In specialized cells,the majority of programs are switched off, allowing them to efficiently focus on a given task. This is what epigenetic mechanisms do: They package and organize the DNA, such that certain bits are shielded away and silenced, while other parts are accessible and readily executable. As such, epigenetic mechanisms are vital for normal development and health. For instance, in the absence of certain epigenetic factors embryonic stem cells fail to differentiate. Epigenetic malfunctioning has also been observed in various diseases: For example, if normally silenced programs become activated, cells may change their identity; white blood cells, for instance, can turn into cancerous cells when their epigenetic machinery is faulty. The epigenome comprises a number of chemical alterations, which exist 'on top' of the DNA sequence itself. For example, at the occurrence of certain DNA sequence features, methyl groups can be added to the DNA to silence corresponding genetic elements. Additionally, the DNA sequence is wrapped around histone proteins forming a "beads-on-a-string" type of architecture. By chemically modifying individual histone proteins, neighboring 'beads' can be brought into tight contact with each other thus forming dense and inaccessible regions of DNA. Alternatively, a different set of Histone modifications can result in open and accessible DNA domains. Histone modifications are dynamically established by a large set of different enzymes, so called 'epigenetic writers'. They can also be actively removed by a number of specific 'epigenomic erasers'. The thus established epigenomic patterns are recognized by 'epigenetic readers'. Interestingly, some steady-state epigenomic modifications are remarkably well correlated with transcriptional activity, suggesting that effector proteins are indeed providing a read-out of epigenomic patterns. These findings have lead to the histone code hypothesis, according to which transcriptional activity is regulated by epigenomic modifications. However, despite intense research and substantial progress in our understanding of epigenetic mechanisms, the histone code has remained enigmatic. Technological advances in the measurement of epigenomic snapshots have led to an explosion of available data. Yet owing to the high complexity and changing nature of these marks, a precise understanding of their meaning and readout is lacking. Today, I see a unique opportunity to tackle this challenge with the help of sophisticated machine learning technologies: These methods use computer systems to 'learn' hidden relationships from large data sets. I will build new computational tools to capture the molecular mechanisms underpinning the dynamic changes of epigenomic marks. Along with my co-investigator, I suggest cycling between sophisticated computational predictions and wet lab experiments that provide dynamic profiles of epigenomic patterns. In particular we plan to disturb the epigenetic machinery by rapidly degrading individual writers to observe how their action orchestrates operations of other writers and readers. I will also use statistical methods to analyse the spatiotemporal correlation between dynamic epigenomes and changing gene expression. This project will benefit from the existing epigenomic expertise at Dundee University and our efforts will in turn inform on-going projects to understand epigenetic contributions to healthy development and disease. In addition, parts of the project will be carried out at the Cyber Valley Campus Tuebingen, which hosts some of the world leaders in causal machine learning techniques.

View the original record at the funder ↗

Researchers

Gabriele Schweikert (Principal Investigator)Tom Owen-Hughes (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Machine Learning Methods for Dynamic Gene Regulation and Personalised Epigenomics
Novel Machine Learning Techniques to Elucidate Function and Dynamics of Epigenomic Mechanisms
Investigating the roles of H3K14 and H3K18 acetylations in controlling gene expression and cell phenotype
Novel approaches for epigenomic profiling of repetitive elements
Machine-learning to create predictive models of genetic regulation

Original classification

Fellowship

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.