InterPro and Pfam: Protein domains and families for biomedical research
In plain English
AI plain-English summaryEvery time a scientist sequences a new genome, a hidden layer of proteins must be decoded to make sense of the DNA. InterPro is the central database that classifies these proteins into families and domains, drawing on 13 separate resources including Pfam. The problem is that the sheer volume of new sequences has exploded, and traditional methods cannot keep up. Meanwhile, AI and machine learning tools now offer a way to extract biological meaning from this flood of data. This project will integrate those approaches—adding KOfams and AMRFinder databases, applying ML methods, and reorganising protein family data—to scale InterPro to handle billions of proteins. If successful, the resource will accelerate biomedical research in concrete ways: helping scientists interpret human mutations, predict drug interactions, track antimicrobial resistance, and respond to emerging pathogens. This is a fundamental infrastructure project. It does not directly treat a disease or build a device, but without it, much of the genomic data generated today would remain uninterpretable—a silent bottleneck in nearly every area of molecular biology.
View original technical description
View the original record at the funder ↗
Researchers
Related Research
Grants with similar aims, by meaning.
Original classification
Biomedical Resources GrantPlain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research. Is something wrong? Let us know