Completed Genetics & Molecular Biology Plants, Animals & Ecology

Deciphering the regulatory genetic code

In plain English

AI plain-English summary

Every human cell carries roughly 10 million DNA spelling differences compared to any other person, and most of them sit in stretches of DNA that do not code for proteins. This project aims to crack the "second genetic code" that governs when and where genes are switched on or off, by studying how proteins called transcription factors read non-coding DNA sequences. The first genetic code—how DNA letters become proteins—was solved decades ago. But the vast majority of human genetic variation falls outside protein-coding regions, and its effects remain largely unknown. Without understanding this regulatory code, scientists cannot explain why two people with identical protein-coding genes can have different lifespans, disease risks, or responses to drugs. This is fundamental science. The researchers will generate large datasets using high-throughput laboratory methods, then analyse them with artificial intelligence tools to infer the rules of gene regulation. If successful, the work will produce new computational methods and a deeper biological understanding. In the longer term, deciphering non-coding variants could improve plant and animal breeding, and help explain human variation in health and development—but immediate practical applications are not the goal.

View original technical description
Our project aims to understand how the cell can read the information that is written in its genome. The project is very much a collaboration between biological and computational scientists, who will work closely together to understand basic mechanisms of how cells can tell when and where the genes written in their DNA should be active. In other words, we seek to understand 'the second genetic code'. The first genetic code that describes how DNA sequence is converted to protein sequence was decoded more than 50 years ago, and the first draft of human genome, which describes the sequence of the chemical letters A, C, G and T found in all human cells was published in 2001. However, just knowing the order of the letters is not enough to understand how they instruct cells to function and to grow. Advances in DNA-sequencing have also allowed sequencing of entire genomes of individual humans and we have learnt for instance that sequences of unrelated individuals are different by approximately 10 million DNA bases. These specific variants, and the mechanisms by which they act are largely unknown. This is because most of the changes do not affect protein structure. Instead, the variations are presumed to affect the amount of proteins made in particular cells, by affecting DNA binding of proteins called transcription factors. Our research project aims to understand how the transcription factors read the genomic code. This will also help us to understand how mutations or variations in DNA sequences change the activity of genes. The work is basic research utilizing novel high throughput methods and artificial intelligence based computational data analysis tools. The project will first generate vast amounts of data in a laboratory, and then utilize and understand it using tailor-made computer programs. The work will have immediate benefits to the scientific community in terms of deeper understanding, novel methods and computational tools. The work will also lead to increased understanding of the function of cells during growth and development. In a wider context, the proposed project is part of a broader effort to use advanced genomic and computational tools to understand the basis of human and animal biology. Genetic variants that are located between genes are so common that most plants, animals and people have many of them. To understand their function will be of great benefit to the scientific community in terms of deeper understanding of biological principles, for development of novel experimental methods and computational tools, and also eventually for applications such as plant and animal breeding. In addition, understanding the effect of non-coding genetic variants will help to explain human variation, for example in lifespan and health.

View the original record at the funder ↗

Researchers

Jussi Taipale (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Molecular mechanisms conferring risk for colorectal cancer
The Arabidopsis Epitranscriptome
Computational prediction and analysis of long non-coding RNAs
Deciphering the Non-Coding Genome using Millions of Diverse Whole-Genome Sequences
Using whole genome sequencing to identify non-coding elements associated with diabetes and related traits across ancestries

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.