Active Genetics & Molecular Biology Bones, Joints & Muscles

Improving the accuracy, functionality, scalability, and usability of orthology inference for biological research

In plain English

AI plain-English summary

Every time a biologist uses a mouse to study a human disease, or a fruit fly to understand a genetic disorder, they rely on a computational tool called OrthoFinder to match equivalent genes across species. That tool is only about 80% accurate on standard tests, and it fails to make any match at all for thousands of genes in a typical analysis. It also struggles to handle the flood of genome data now being produced, and its command-line interface locks out researchers who lack programming skills. This project will overhaul OrthoFinder to fix each of those weaknesses: boosting its accuracy, expanding the number of genes it can analyse, scaling it to handle current and future genome datasets, and building a graphical interface so any biologist can use it without typing code. If successful, the improvements will ripple through thousands of studies that depend on OrthoFinder to transfer knowledge between species, and will also raise the quality of public sequence databases that rely on the method. The work is squarely aimed at improving research infrastructure rather than producing immediate clinical or commercial applications, but better orthology inference directly strengthens the foundation of comparative genomics, model-organism research, and evolutionary biology.

View original technical description
Inferring the phylogenetic relationships between biological sequences (orthology inference) is fundamental to biological and biomedical research. It provides the framework for transfer of biological knowledge between species, and enables the use of model organisms for studying health and disease. However, orthology inference methods are not perfect. The best methods are only ~80% accurate on benchmark tests and also fail to make inferences for thousands of genes in a typical analysis. In addition to these limitations, orthology inference methods are poorly scalable and are unable to analyze the current (or future) quantities of genome data. Finally, the methods themselves are poorly accessible to researchers lacking expertise in command line environments. This project aims to address each of these limitations and challenges by improving the accuracy, enhancing the functionality, increasing the scalability, and extending the usability of OrthoFinder. By delivering these improvements this project will have impact on the accuracy and capability of thousands of studies that are underpinned by OrthoFinder. It will also improve the accuracy and utility of repositories of biological sequence data that rely on the method. Finally, it will improve access to high-level comparative genomics tools within the biological and biomedical sciences.

View the original record at the funder ↗

Researchers

Steven Kelly (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

14 NSFBIO:Towards detailed and consistent function prediction from protein family databases
18-BBSRC-NSF/BIO : CIBR:Implementing an explicit phylogenetic framework for large-scale protein sequence annotation
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
Further development of the PSIPRED server into an integrated tool for systems biology and functional genomics researchers
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam

Original classification

Discovery Award

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.