2024BBSRC-NSF/BIO: Development of rich AI datasets for furthering life science research
In plain English
AI plain-English summaryEvery day, researchers around the world manually type protein structure data into a global archive—a slow, error-prone process this project aims to replace with automation. The Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) hold the three-dimensional shapes of biological molecules, which are essential for understanding how cells work and for designing drugs. Currently, scientists must manually submit each structure, a laborious step that introduces mistakes and slows down the release of new data. This project will build software tools that let researchers deposit structures directly from the four major data-analysis programs (Phenix, CCP4, CCP-EM, and Global Phasing), eliminating manual entry. It will also enable the batch submission of multiple related structures at once, grouping them into coherent investigations. If successful, the quality and completeness of the PDB and EMDB archives will improve significantly. Researchers and educators across the life sciences will have faster access to cleaner, better-organised structural data. This is a fundamental infrastructure project—it does not directly cure a disease or engineer a new material, but it quietly underpins nearly every discovery that does.
View original technical description
View the original record at the funder ↗
Researchers
Related Research
Grants with similar aims, by meaning.
Original classification
Research GrantPlain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research. Is something wrong? Let us know