Active Genetics & Molecular Biology Plants, Animals & Ecology

What determines protein abundance in plants?

In plain English

AI plain-English summary

A single plant’s cells can contain thousands of different proteins, and scientists still do not fully understand why some proteins are abundant while others are scarce. This project uses the model plant *Arabidopsis thaliana* to untangle the multiple, interacting mechanisms that control protein levels—from DNA structure and chemical marks to RNA production and protein breakdown. Most research has focused on gene transcription because it is easy to measure, but transcription alone is a poor predictor of how much protein actually ends up in a cell. Translation and protein degradation matter greatly, especially in plants, yet they remain understudied. The team will combine genome-wide measurements with advanced statistical methods to apportion the relative contributions of each regulatory step, using a genetically diverse population of *Arabidopsis* lines. This is fundamental science: it asks how a complex organism translates its genetic blueprint into functional molecules. If successful, the work will produce a public, user-friendly resource for gene mining and provide a systems-level understanding of protein regulation. That knowledge could eventually guide efforts to engineer crops with higher yields, better stress tolerance, or improved nutritional content—but the immediate payoff is a clearer picture of a core biological process.

View original technical description
Proteins are the workhorses of the cell: they facilitate chemical reactions, act as gene switches and have structural roles. For cells to work efficiently, proteins need to be produced in the right place, at the right time and in the right amount. They also need to be removed when no longer needed. Crick's Central Dogma states that coding sequences of DNA are transcribed into mRNAs, which in turn are translated into proteins. There are many levels at which this process is regulated and there are still many gaps in our knowledge. We expect both inherited and environmental differences between individuals to play important roles in the control of proteins. This project seeks to use the model plant, Arabidopsis thaliana, to answer fundamental questions about the control of protein expression, including which mechanisms are important and how they interact in a complex multi-cellular organism. We also aim to determine to what extent the protein content of a given cell, tissue or organ predicts observable traits (the phenotype) of the plant. To address these questions, we have designed an integrated programme of experiments and sophisticated mathematical analysis around a genetically variable population of Arabidopsis (known as the MAGIC population). This is a powerful genetic resource for mapping sections of DNA that correlate with variation in a trait (known as quantitative trait loci, QTL), to identify causal variants and dissect the regulation of genome expression. We will characterise and compare the following different processes that potentially influence protein expression in the MAGIC lines: 1. Structural variation within the genome (including small-scale variation and large-scale structural rearrangements) 2. Chromatin accessibility, a measure of the availability of a given region of DNA for transcription. 3. Chemical modifications to DNA that do not involve a change in DNA sequence, known as epigenetic marks, which often indicate environmental perturbation. 3. mRNA abundance. 4. Protein abundance. It is important to take an holistic approach, because the amount of any given protein in an individual is determined by the balance of these processes. Much effort has been spent studying gene transcription, because it is relatively easy to measure on a genome-wide scale. However, evidence suggests that transcription is a poor predictor of protein abundance, because the control of translation and protein degradation are important, particularly in plants. Less research has been done on measuring translation, protein amount and protein breakdown but advances in technology now let us do so. Although it is relatively straightforward to measure genomic structural variation and epigenetic marks such as DNA methylation, their impact on protein expression is unclear. Therefore, we are in an exciting position to provide enormous insight into protein regulation. The power of this project derives from innovative computational analysis that will enable us to apportion the relative contributions of genotype, transcription, protein synthesis and protein degradation and identify networks controlling protein expression. Because collecting genome-scale data from many samples is expensive and time-consuming, we will use novel statistical methods to get more information without significantly increasing sample size, including combining different layers of information. This will be the first study of this kind on this scale. As well as depositing our data in public repositories, our findings will be made available to the academic community via a user-friendly knowledge discovery and gene mining resource. The approaches developed in this project will provide valuable fundamental insights that will be applicable to other organisms and which will also pave the way to future crop improvement.

View the original record at the funder ↗

Researchers

Frederica Theodoulou (Principal Investigator)Gancho Slavov (Co-Investigator)Kathryn Lilley (Co-Investigator)Keywan Hassani-Pak (Co-Investigator)Richard Mott (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

The Arabidopsis Epitranscriptome
Defining the molecular basis of chloroplast transcription of photosynthetic genes
The non-coding Arabidopsis genome
Structure and function of the chloroplast transcription machinery
SILAC proteomics for quantitation of protein isoforms from alternative splicing in Arabidopsis seedlings

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.