Active Chemistry Cells, Biochemistry & Physiology

SureChEMBL: open patent data for all

In plain English

AI plain-English summary

Every month, around 80,000 newly patented chemical compounds are extracted and stored in SureChEMBL, a free, open-access database that already holds more than 20 million unique chemical structures from roughly 50 million patents. This matters because most patent data is locked behind paywalls or scattered across incompatible formats, making it difficult for researchers to see what molecules have already been patented before starting their own drug development work. Without open access, scientists waste time and money rediscovering known compounds or inadvertently infringing on existing patents. If this upgrade succeeds, the improved SureChEMBL will let researchers search not just by chemical structure but also by biological targets—such as proteins or genes—and assess how relevant a patent is to their question. A new application programming interface (API) will allow computer programs to query the database automatically, enabling large-scale data mining. The result would be a faster, more reliable way for the global life-sciences community to navigate the patent landscape, reducing duplication of effort and accelerating the early stages of drug discovery. This is infrastructure research: it does not create a new drug itself, but it quietly underpins the entire process of finding one.

View original technical description
SureChEMBL is a large-scale, fully automated, chemical-structure-enabled database providing the research community with open, free and FAIR access to the patent literature. SureChEMBL contains ~140 million patents with ~50,000 added monthly. Of these, ~50 million patents are chemically annotated with more than 20 million unique chemical structures. Around 80,000 new compounds are extracted and stored monthly. The data in SureChEMBL can currently be accessed via a web interface that enables users to perform text and chemical structure queries, filter the output and then display the results. The complete set of chemical structures and patent associations are also available for download. In this proposal we plan to significantly enhance SureChEMBL, enabling a wider selection of questions and use cases to be addressed, by: expanding annotations to include biological entities; developing methods to assess the relevance of information identified in patents and enable more accurate data retrieval and analysis; significantly improving the technical infrastructure and web interface with enhanced stability and usability; providing programmatic access through development of a RESTful application programming interface (API). The new, upgraded SureChEMBL will provide the broader life sciences research community with unparalleled access to a large and rich source of information and knowledge.

View the original record at the funder ↗

Researchers

Andrew Leach (EPMC Awardee)Barbara Zdrazil (EPMC Awardee)

Related Research

Grants with similar aims, by meaning.

The ChEMBL Database An Open Resource for Drug Discovery
The ChEMBL database
BioChemGRAPH - an integrated knowledge graph to facilitate basic and translational research
Re-engineering ChEBI for a sustainable future
BBSRC-NSF/BIO - Expanding fold library in the twilight zone to facilitate structure determination of macromolecular machines

Original classification

Biomedical Resources Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.