Active Computing & AI

2024BBSRC-NSF/BIO: Development of rich AI datasets for furthering life science research

In plain English

AI plain-English summary

Every day, researchers around the world manually type protein structure data into a global archive—a slow, error-prone process this project aims to replace with automation. The Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) hold the three-dimensional shapes of biological molecules, which are essential for understanding how cells work and for designing drugs. Currently, scientists must manually submit each structure, a laborious step that introduces mistakes and slows down the release of new data. This project will build software tools that let researchers deposit structures directly from the four major data-analysis programs (Phenix, CCP4, CCP-EM, and Global Phasing), eliminating manual entry. It will also enable the batch submission of multiple related structures at once, grouping them into coherent investigations. If successful, the quality and completeness of the PDB and EMDB archives will improve significantly. Researchers and educators across the life sciences will have faster access to cleaner, better-organised structural data. This is a fundamental infrastructure project—it does not directly cure a disease or engineer a new material, but it quietly underpins nearly every discovery that does.

View original technical description
The vision of this US RCSB Protein Data Bank/Protein Data Bank in Europe project is to automate data deposition and improve management of three-dimensional (3D) biostructure information stored in the Protein Data Bank (PDB) and the Electron Microscopy Data Bank (EMDB). 3D structures stored in the PDB archive provide a wealth of insights into biochemical and cellular function of biological macromolecules. Historically, each structure determination experiment resulted in a single 3D structure, which was deposited manually to the PDB and identified with a persistent identifier or PDB ID (e.g.,1vol). Primary experimental data underpinning PDB structures are stored in the PDB (macromolecular crystallography or MX: diffraction data) and EMDB (3D electron microscopy of 3DEM: electric Coulomb potential maps). Working with the four major MX and 3DEM structures determination software providers (Phenix, CCP4, CCP-EM, and Global Phasing), we will enable automated one-at-a-time deposition of individual MX and 3DEM structures directly from their programs, thereby avoiding laborious, error-prone manual data entry. We will also enable automated parallel deposition of multiple related MX and 3DEM structures, constituting an investigation. This work will benefit researchers and educators/students across the sciences by improving 3D biostructure data quality and completeness and grouping related structures, which together explain important biological phenomena.

View the original record at the funder ↗

Researchers

Eugene Krissinel (Co-Investigator)Kyle Morris (Co-Investigator)Sameer Velankar (Principal Investigator)Tom Burnley (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

CIBR 19-BBSRC-NSF/BIO: Next generation PDB - FACT infrastructure with value added FAIR data supporting diverse research and education user communities
BBSRC-NSF/BIO - Expanding fold library in the twilight zone to facilitate structure determination of macromolecular machines
Supporting archival and dissemination of small-angle scattering data for atomistic structures in the PDB
Protein Data Bank in Europe - an integrated resource for 3D molecular and cellular structure.
24BBR: Enabling efficient basic and translational research through sustainable expansion of PDBe Knowledge Base Using Integrative Resource Framework

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.