Active Computing & AI Plants, Animals & Ecology

Data as a foundation for AI innovation and global discovery research in the life sciences

In plain English

AI plain-English summary

EMBL-EBI’s data resources—the world’s most-used collection of biomolecular information—are being rebuilt to feed artificial intelligence and connect scientists in low- and middle-income countries (LMICs) to the global research ecosystem. The problem is that biological data has grown too vast, too scattered, and too poorly formatted for modern AI tools to exploit. Researchers in LMICs often cannot access or contribute to these resources, creating a lopsided global data system that limits progress on shared challenges like human health and biodiversity loss. If this succeeds, AI developers will get ready-to-use training datasets stripped of formatting barriers, while EMBL-EBI’s own data services will become faster and more intelligent. Standardised, linked data will let software developers and researchers find and reuse information across different fields without manual wrangling. For LMICs, the shift means their scientists can both contribute data and benefit from discoveries made elsewhere—building an equitable global infrastructure that underpins everything from drug discovery to crop resilience. This is fundamental infrastructure work: invisible to most, but essential for the next generation of life-science breakthroughs.

View original technical description
EMBL-EBI's data resources are a fundamental necessity for biomolecular research. Through open data, there is tremendous potential for accelerating scientific understanding and producing significant gains for all nations. The transformative scientific frontiers in which EMBL-EBI can play a unique role are the development of new AI algorithms, the integration of multimodal data, and a deeper engagement with LMICs. Alongside meeting the ongoing challenges of managing data scale and diversity, we will co-develop new interfaces to data delivery. Firstly, for AI, we will deliver cross-cutting data training sets, enriched and formatted to lower barriers for innovators. Deploying AI in our processes will increase efficiencies as well as bring the fruits of AI to users in our data services. Secondly, data resources need to interoperate consistently to ease the submission of FAIR, linked, multimodal data and maximise the downstream reuse potential for software developers and researchers via easy-to-use interfaces to find and explore data. Finally, the emphasis on LMICs within our engagement programme is essential to build an equitable and interoperable global data ecosystem; without this, our ability to impact on global challenge areas such as human health and biodiversity will be severely diminished.

View the original record at the funder ↗

Researchers

Ewan Birney (EPMC Awardee)Helen Parkinson (EPMC Awardee)Jo McEntyre (EPMC Awardee)Robert Finn (EPMC Awardee)Sameer Velankar (EPMC Awardee)Thomas Keane (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

Building a Next Generation Image Repository: Molecular Annotation and Cloud-based Data Processing and Analysis
2024BBSRC-NSF/BIO: Development of rich AI datasets for furthering life science research
EMDB: dealing with the cryo-EM data deluge
EMBL-EBI expansion: response to the exponential increase in life science data
The Image Data Resource: Making Biological Imaging Data FAIR

Original classification

Discretionary Award

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.