Completed Genetics & Molecular Biology Cells, Biochemistry & Physiology

InterPro and Pfam: Protein domains and families for biomedical research

In plain English

AI plain-English summary

Every time a scientist sequences a new genome, a hidden layer of proteins must be decoded to make sense of the DNA. InterPro is the central database that classifies these proteins into families and domains, drawing on 13 separate resources including Pfam. The problem is that the sheer volume of new sequences has exploded, and traditional methods cannot keep up. Meanwhile, AI and machine learning tools now offer a way to extract biological meaning from this flood of data. This project will integrate those approaches—adding KOfams and AMRFinder databases, applying ML methods, and reorganising protein family data—to scale InterPro to handle billions of proteins. If successful, the resource will accelerate biomedical research in concrete ways: helping scientists interpret human mutations, predict drug interactions, track antimicrobial resistance, and respond to emerging pathogens. This is a fundamental infrastructure project. It does not directly treat a disease or build a device, but without it, much of the genomic data generated today would remain uninterpretable—a silent bottleneck in nearly every area of molecular biology.

View original technical description
InterPro is a core protein data resource that amalgamates 13 protein family databases (including Pfam) to provide the definitive description and classification of protein domains and families. Various scientific communities rely on the range of annotations provided by InterPro and its member databases to acquire novel insights into the vast amounts of new DNA sequence data, enabling scientists to interpret experiments and design new ones based on annotations spanning complete genomes down to single residues. The past five years have witnessed a major increase in the scale of sequences, concomitant with the emergence of AI/ML approaches for unlocking the biological signals encoded within the data. This project will exploit these innovative approaches to enhance protein annotations and refine classifications through the incorporation of additional data types, along with the application of ML methods and organisation of the protein family data. We will integrate KOfams and AMRFinder into InterPro, while continuing to scale the InterPro/Pfam/HMMER resources to tackle billions of proteins. Cumulatively, these developments will accelerate biomedical research by facilitating scientists to decipher the effect of human mutations, understand drug interactions, combat antimicrobial resistance, and tackle emerging pathogens.

View the original record at the funder ↗

Researchers

Alex Bateman (EPMC Awardee)Robert Finn (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
Keeping pace with protein sequence annotation; consolidating and enhancing Pfam and InterPro's methodologies for functional prediction
IDA2GO - Improving Domain Annotation and Representation within InterPro
UniFam: an integrated protein families resource .
14 NSFBIO:Towards detailed and consistent function prediction from protein family databases

Original classification

Biomedical Resources Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.