Wheat, barley, rice, and other staple crops now have their full genomes sequenced, but the data sits in scattered formats that breeders cannot easily use to find the genes behind better yields or disease resistance. The problem is that knowing a crop’s genome sequence is only the first step. To breed more nutritious or resilient varieties, researchers need to compare hundreds of individuals from the same species, linking genetic variants to observable traits like drought tolerance or grain size. Currently, no unified platform lets them do this across multiple crops using publicly funded data. This project will build that platform—an open-source, web-accessible system that stores genetic variation data alongside plant phenotype records, and provides tools to query both together. If successful, the platform will lower the barrier for academic and industrial breeding programmes to use genomic information directly. Instead of each group building its own database, a shared infrastructure will let researchers quickly identify which genetic markers are associated with a desired trait, accelerating the development of more productive crop varieties. The platform will also be packaged as a virtual machine for local installation, making it usable by labs without high-performance computing access. This is primarily an infrastructure project—it does not discover new genes itself, but creates the tools that make such discoveries routine.
View original technical description
Recent advances in sequencing technologies and computational tools have made it possible to sequence the genomes of some of the world's most important crop species, such as rice, barley, rapeseed, maize, soya and wheat. These crops constitute a substantial part of the daily food intake for most of the population of the world and any improvements in the breeding for more efficient and nutritious varieties will have a direct impact on ensuring global food security. Whilst obtaining the genome sequences for these crops provides a hugely useful resource for giving insights into the differences between species, it is through sequencing different individuals from the same or closely-related species which allows us to identify useful genetic variants which can be selected for during plant breeding. These approaches require a combination of sequence and phenotypic data, plus analysis tools. We propose to develop a crop bioinformatics platform which enables users to access this genetic and phenotypic variation and perform analyses to explore gene expression and associations between genetic variation and traits. The platform will be developed using open source principles and publicly available data. Population-wide genetic variants will be represented on a genomic data structure; an archiving system for storing plant phenotype data will be developed; tools to allow the querying of these datasets and analyses to link genotype to phenotype will be implemented; and the platform will be accessible via TGAC and EBI servers but also packaged into a virtual machine for easy installation on users' local hardware. This novel platform for crop bioinformatics will promote opportunities for collaborative work with R&D groups in industry, research and academia. The availability of data generated by publicly funded resources, and the concomitant development of new, production-quality tools will lower the barriers to information-enabled crop science, stimulating new opportunities for research and application. The platform will also open up new opportunities for the UK bioinformatics community, traditionally focused on biomedical applications, by developing alternative career paths around biotechnology and agri-food.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know