Active Genetics & Molecular Biology Pregnancy, Children & Inherited Conditions

Determining the role of untranslated regions of the genome in disease

In plain English

AI plain-English summary

Only 1% of human DNA directly codes for proteins, but most disease-linked genetic variants lie in the other 99%—and this project will systematically map how variants in the "untranslated regions" at the ends of genes alter protein production and cause illness. Current genetic testing mostly ignores these regions because their role was too expensive to study at scale. The team will analyse DNA and medical data from over one million people, including UK Biobank participants, to link variants in untranslated regions to both common diseases like diabetes and rare developmental disorders. They are building a computational tool that predicts whether a given variant changes how much protein a gene makes, then testing those predictions against real protein levels in blood. If successful, the tool could give diagnoses to families with rare diseases that have remained unexplained for years, and identify new drug targets for common conditions by revealing which proteins are over- or under-produced. The researchers will release the tool publicly, allowing other scientists to probe untranslated regions without starting from scratch. This is fundamental science with a direct diagnostic payoff—it fills a blind spot in how we interpret the human genome.

View original technical description
Context: New technologies have recently made it possible to explore how differences across millions of people's DNA (called "variants") can affect the production of proteins and contribute to disease. DNA provides our cells with instructions for making proteins, and variants can alter the type or amount of protein produced, potentially causing disease. Although many variants have no effect, some can increase the risk of common diseases, whilst others can cause very rare diseases. Tens of thousands of variants have already been linked with thousands of different diseases. However, we still don't know why most variants predispose to common disease, and more than half of rare diseases remain undiagnosed. Challenge: Only around 1% of our DNA codes for genes that provide instructions for making proteins. Our understanding of why genetic variants cause disease is mostly limited to this 1%, but most variants linked to diseases occur in the other 99%. Here, we will focus on a specific part of DNA, called "untranslated regions," which do not provide instructions for making proteins but sit at both ends of protein-coding genes. These regions play a key role in controlling how much protein is made from each gene. Because variants in these regions could directly affect protein levels, their impact is easier to measure than variants in other parts of DNA. However, the effect of variants in these untranslated regions has been largely unexplored, as sequencing DNA in large numbers of people was previously too expensive. This project will use large recently-generated datasets to find new links between disease and variants in the untranslated regions of genes. Aims and Objectives: We will use genetic and medical data on >1 million people to understand how differences in people's DNA can affect the production of proteins and contribute to disease. We have already found that certain variants in untranslated regions can alter protein levels and lead to disease. Now, we want to expand our research to look at the effect of all genetic variants across untranslated regions. We will: Develop a new tool to better predict the effects of genetic variants in untranslated regions (for example, whether they are likely to change how much protein is made, and by how much). Test our tool using data on levels of proteins circulating in blood (using data from a large number of individuals who are part of UK Biobank). Use our tool to see if variants in untranslated regions are linked to hundreds of common diseases (for example, diabetes or heart disease). Apply our tool to find new diagnoses in untranslated regions for patients with rare diseases (for example, developmental disorders). Share our tool to enable other researchers to investigate untranslated regions. Applications and Benefits: New links between genetic variants and disease could provide diagnoses for families affected by rare diseases and help identify new drug targets that treat diseases by adjusting protein levels. By making our tool publicly available for other researchers to use, our work will provide new insights into the important role that untranslated regions play in health and disease. Our research is important for understanding how diseases work, improving diagnosis, and designing new treatments.

View the original record at the funder ↗

Researchers

Caroline Wright (Principal Investigator)Gareth Hawkes (Co-Investigator)Leigh Jackson (Co-Investigator)Michael Weedon (Co-Investigator)Nicky Whiffin (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Human functional genomics of post-translationally modifying clinical coding variants: FGx-PTMv
Identification of novel short tandem repeat expansions in neurological disorders
Computational analysis of protein covariation for the identification of disease-associated variants in coding regions
Understanding the impact of ‘near-coding’ variation in human disease
Translational genomics- maximising potential for NHS patient care.

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.