Completed Cells, Biochemistry & Physiology Genetics & Molecular Biology

The PRIDE database: A proteomics data hub in the life sciences

In plain English

AI plain-English summary

Every time a scientist runs a mass spectrometer to identify the proteins in a blood sample, the resulting data pile can be enormous—and PRIDE is the central digital warehouse where that data goes to be stored, shared, and reused. Established in 2004 at EMBL-EBI, PRIDE is the world’s leading proteomics data repository and, since 2011, has coordinated the ProteomeXchange Consortium, which standardises how researchers worldwide submit and share protein data. The problem it addresses is simple: as clinical studies generate ever-larger datasets, the old manual submission methods are becoming a bottleneck. This project will build an automated submission interface that third-party software can plug into directly, and create open quality-control pipelines to ensure only reliable data flows onward to major biological resources like UniProt and Ensembl. It will also develop reproducible analysis pipelines tailored for multi-omics approaches—combining protein data with genomics or metabolomics—which are increasingly used in personalised medicine. If successful, PRIDE will quietly underpin the infrastructure that lets any biologist, not just proteomics specialists, reuse high-quality protein data to understand disease mechanisms or identify drug targets.

View original technical description
Established in 2004 at EMBL-EBI, PRIDE is the world-leading proteomics data repository and, since 2011, is leading the ProteomeXchange Consortium, standardising public proteomics data submission and dissemination worldwide. The success of PRIDE and ProteomeXchange, has largely driven the proteomics community to widely embracing open data policies. To continue serving and shaping this increasingly prominent field, we primarily request support for the further development of PRIDE as a repository, to enable an efficient handling of proteomics ‘big data’ such as the increasingly large datasets generated from clinical studies. A key component is the development of a submission Application Programming Interface that can be integrated in third-party software. Second, PRIDE will become a Hub for proteomics data, by establishing robust data dissemination pipelines to key resources (UniProt, Ensembl/Ensembl Genomes, Expression Atlas), enabling proteomics data reuse by all biological researchers. We will build novel, open quality control pipelines to ensure that only high-quality data is propagated. Third, also to facilitate data reuse, we will develop and make available open, reproducible proteomics data analysis pipelines tailored for multi-omics approaches used e.g. in personalised medicine. These pipelines will be connected to PRIDE, bringing cutting-edge analysis tools closer to the data, as datasets become larger.

View the original record at the funder ↗

Researchers

Juan Antonio Vizcaino (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Extending the PRIDE database ecosystem: PRIDE-Controlled Access (PRIDE-CA) and the PRIDE-Knowledgebase (PRIDE-KB)
The Proteomics Identifications Database (PRIDE).
GRAPPA - Global compRehensive Atlas of Peptide and Protein Abundance
3D-Proteomics: FAIRification of proteomics data for comprehensive integration with structural biology information
PROCESS - Proteomics data Collection, Software and Standards to support open access and long term management of data

Original classification

Biomedical Resources Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.