Proteins are built from modular building blocks called domains, and this project will train artificial intelligence to read those domains like sentences to predict what proteins do and which other proteins they stick to. Current protein language models treat whole protein sequences the way a language model might treat an entire novel. This project argues that focusing on domains—the discrete, functional paragraphs within that novel—will produce sharper predictions. The team will use the Encyclopedia of Domains, a curated database of these units, as training data. The core gap is that existing models often miss how domains interact with each other, which is what actually drives protein behaviour inside cells. If the models work, researchers could predict protein-protein interactions more accurately. That matters because those interactions control nearly every process in a cell—signalling, metabolism, replication. Better predictions would speed up drug target identification and help design synthetic proteins for industrial enzymes or agricultural crops. The work is fundamental computational biology; it will not directly treat a disease tomorrow. But similar fundamental work on sequence-based language models has already transformed how biologists search for new antibiotics and understand genetic mutations.
View original technical description
This proposal seeks to further expand protein language modeling by developing next-generation models using a recently published resource, the Encyclopedia of Domains (TED), as the main foundation for training data. Protein language models, much like those used for natural languages such as English, learn to recognize patterns, structures, and relationships within sequences—in this case, sequences of amino acids that make up proteins. The focus of this project is on protein domains, which are distinct structural and functional units within proteins. Domains play a crucial role in defining how proteins fold, interact, and carry out their biological functions. Beyond simply identifying these domains, the project will also emphasize understanding the interactions between domains, as these interactions often dictate how proteins interact with one another or with other molecules in a cell. By concentrating on domains and their interactions, the project aims to create protein language models that are not only more precise but also more versatile. Such models could transform our ability to predict protein-protein interactions (PPIs)—an essential area of research, as these interactions underlie almost every cellular process. In addition, the models will be tailored to improve our understanding of protein function, a key piece in the puzzle of decoding the human proteome and other complex biological systems. This work has the potential to revolutionize fields like biotechnology, molecular biology, and drug discovery. For example, better predictions of protein interactions can accelerate the identification of drug targets or the design of synthetic proteins with specific functions. Moreover, these enhanced models could contribute to solving pressing challenges, such as developing treatments for diseases caused by misfolded or malfunctioning proteins. By leveraging the Encyclopedia of Domains, a highly curated and detailed dataset, the project builds upon a rich resource that provides insights into the architecture and behavior of proteins. Integrating this data into advanced computational models represents a significant step forward in bridging the gap between computational biology and real-world applications. This initiative would not only deepen our understanding of protein biology, but also could provide groundwork for future innovations in health, agriculture, and industrial applications.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know