Completed Chemistry Cells, Biochemistry & Physiology

The ChEMBL Database An Open Resource for Drug Discovery

In plain English

AI plain-English summary

Drug hunters can now search over 12 million experimental bioactivity measurements on 1.3 million distinct compounds, all free to use. This matters because the drug discovery sector has been shrinking. Large commercial firms have downsized early-stage R&D, and new entrants lack the high-quality open data needed to design and optimise new compounds. The ChEMBL database fills that gap by capturing published bioactivity and structure-activity relationship data, adding curation and semantic annotation so researchers can data-mine and search it effectively. If this work succeeds, the database will expand in six directions: deeper coverage of bioactivity space, better indexing with ontologies, inclusion of patent literature, annotation of resistance and natural population variation, technology upgrades such as RDF services and an API, and a broader user community reaching into drug metabolism, pharmacokinetics, clinical research, and biotechnology. The practical impact is on the infrastructure of drug discovery itself—making the early stages cheaper, faster, and more accessible to academic labs and small companies that cannot afford proprietary databases. This is applied, not fundamental science: the goal is to accelerate the translation of genomic and omics data into real treatments.

View original technical description
The translation of sequence, population and omics data into healthcare advances has been frustratingly slow. This has led to many large commercial organisations downsizing their early-stage R&D. A key factor with new entrants to the drug discovery sector is the lack of high quality open data to support compound design and optimisation. This application builds upon a previous Wellcome Trust Strategic Award (WT086151/Z/08/Z) to extend ChEMBL. ChEMBL captures published bioactivity and structure act ivity relationship (SAR) data and adds curation and semantic annotation to allow data-mining and searching. ChEMBL contains 1.3 million distinct compounds and over 12 million experimental bioactivities. In this current application we will develop: 1) Greater coverage of bioactivity space - to deepen and formalise the data contained in ChEMBL. 2) Enhanced indexing with ontologies - to provide more structured data and ease further integration. 3) Patent coverage - extend chemical-structure /target data to include patent literature. 4) Address variation data to include annotation of resistance and natural population variation. 5) Technology enhancements - including RDF services and an API to ease data entry and curation. 6) Expanded user community of ChEMBL - with new beneficiaries in drug metabolism and pharmacokinetics, clinical, and biotechnology communities.

View the original record at the funder ↗

Researchers

John Overington (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

The ChEMBL database
SureChEMBL: open patent data for all
BioChemGRAPH - an integrated knowledge graph to facilitate basic and translational research
Chemogenomics.
Continued development of the ChEBI database and ontology for improved interoperability with biomedical resources

Original classification

Strategic Award - Science

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.