Skip to main content

Posts

Showing posts with the label SureChEMBL

SureChEMBL: Your next source of research data

Slides and recordings from the recent ChEMBL UGM are starting to appear on the meeting website. Here I want to draw attention to the presentation by Nicolas Bosc and myself on " SureChEMBL: Your next source of research data ". Ever since my NextMove Software days, I've been aware of the large amount of scientific data available in patents. This includes everything from large sets of chemical analogs, to bioactivity values, reactions, and NMR spectra. US patents in particular are a rich source of data as they are (a) born digital, and (b) freely available, and thus automated tools to extract relevant data can generate substantial high quality datasets. For example, here's a graph I did back in 2017 to illustrate a NextMove Software blog post entitled "Are more bioactivities available from patents than from the academic literature?". This compared the data deposited from papers into ChEMBL and that extractable by LeadMine from patents (see the talk linked fr...

Recording: SureChEMBL2.0 - Now with Added Disease and Protein Annotations

  A week ago, I had the pleasure of presenting SureChEMBL2.0 at the  Cambridge Cheminformatics Network Meeting , organised by Andreas Bender and kindly hosted by the  Cambridge Crystallographic Data Centre . It was a great opportunity to introduce one of the latest freely available databases of scientifically annotated patents to a broad scientific audience. The recording of the talk is now available  online , along with the  slides . What did I cover during this 30-minute talk? Why scientists should pay attention to patent data Why patents are challenging to work with What SureChEMBL is and what it does How we identify chemical compounds in patent documents What SureChEMBL 2.0 has recently introduced How we annotate patents for genes/proteins and diseases How we are improving the quality of structures extracted from images What you can download from the SureChEMBL core datasets — and what they contain Examples of queries that SureChEMBL h...

Adding Biomedical Annotation to SureChEMBL: Beyond the Chemical Space

        Dear users,   Since its introduction in 2015, SureChEMBL has been a database focused on chemical annotations. We extract compound structures from patent texts, images, and Molfiles when available, and register them in our database . This chemistry-first approach is even reflected in our name.   However, we know that intellectual property documents capture far more than chemistry. This was illustrated by Stefan Senger in 2017 ( 10.1186/s13321-017-0214-2 ), who showed that compound–target interactions can appear years before being mentioned in the scientific literature.   Our first step into biomedical annotation A few years ago, we took a first step beyond chemistry by adding annotations for genes/proteins, diseases, and mechanisms of action in the SureChEMBL UI. These were generated by an in-house Natural Language Processing (NLP) model that performed reasonably well for an initial version.   Example of biomedical annotation in a pa...

Download the SureChEMBL2.0 data: major update!

    Dear SureChEMBL users, This is our second blog post from our series related to SureChEMBL2.0. If you missed the first one , you should take 5 minutes to read it. We know that for many of you, not being able to download new patent data from SureChEMBL for several months has had a significant impact on your work. Please rest assured that halting the data release was not a decision taken lightly. Continuing to develop on legacy software had reached a breaking point and could no longer be sustained within our current resource constraints.  In terms of SureChEMBL downloads, we were previously offering three types, all focused on patent compounds but differing in update frequency and content: MAP files: TSV files released quarterly, listing all compound–patent relationships identified in that period, along with the locations of the compounds within the patents.

 Compound data dump: SDF or TXT files, also released quarterly, containing only the compounds found in that perio...