Completed Genetics & Molecular Biology Cells, Biochemistry & Physiology

Extending the PRIDE database ecosystem: PRIDE-Controlled Access (PRIDE-CA) and the PRIDE-Knowledgebase (PRIDE-KB)

In plain English

AI plain-English summary

The PRIDE database, which stores roughly 83% of the world's publicly shared proteomics data, is being rebuilt to handle sensitive human clinical samples and to make protein-level genetic variation searchable. This matters because proteomics data—the study of proteins expressed by genes—is increasingly used in personalised medicine, but the current system cannot properly handle sensitive patient data or link protein variants to genetic information. Without controlled access, clinical datasets cannot be shared ethically, limiting research into how protein changes drive disease. If successful, this project will create two new resources. PRIDE-Controlled Access (PRIDE-CA) will let researchers securely store and share clinical proteomics data, similar to how DNA sequencing data is already managed. PRIDE-Knowledgebase (PRIDE-KB) will store high-quality re-analyses of proteomics data, starting with proteogenomics—linking protein variants, isoforms, and novel coding events to DNA-level genetic variation. This would allow researchers to search for protein-level genetic changes in the same way they currently search for DNA mutations, using the Beacon framework adapted for proteins. The result is infrastructure that quietly underpins future diagnostic tools and targeted therapies, without which personalised proteomics approaches cannot scale.

View original technical description
The PRIDE database, established in 2004 at the European Bioinformatics Institute (EMBL-EBI), is the world-leading proteomics data repository. PRIDE is also the leading partner in the ProteomeXchange Consortium, standardising public proteomics data submission and dissemination worldwide. PRIDE stores ~83% of all ProteomeXchange datasets, with ~5,300 datasets submitted during 2020 alone. We request support to extend PRIDE, enabling the appropriate handling of clinical sensitive human datasets, through implementing controlled-access (CA) data capabilities. We will leverage infrastructure implemented for CA DNA/RNA sequencing data in the European Genotype Archive at EMBL-EBI, to create PRIDE-CA. Additionally, we will adapt the Beacon framework (developed for genetic variation at the DNA level) to support genetic variation data at the protein level, which is increasingly relevant in personalised medicine studies. We will also develop the PRIDE Knowledge-Base (PRIDE-KB), a complementary resource to PRIDE, to store and disseminate high-quality proteomics data re-analyses, to generate new biological insight. We will implement proteogenomics data as the first use case supported in PRIDE-KB, reporting protein variants, isoforms and novel coding events for personalised proteomics approaches. To support this, diverse components will be developed and integrated, including new data archiving infrastructure, web interface, open data analysis pipelines and data inclusion guidelines.

View the original record at the funder ↗

Researchers

Juan Antonio Vizcaino (EPMC Awardee)Jyoti Choudhary (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

The PRIDE database: A proteomics data hub in the life sciences
The Proteomics Identifications Database (PRIDE).
GRAPPA - Global compRehensive Atlas of Peptide and Protein Abundance
3D-Proteomics: FAIRification of proteomics data for comprehensive integration with structural biology information
PROCESS - Proteomics data Collection, Software and Standards to support open access and long term management of data

Original classification

Biomedical Resources Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.