Skip to main content

Posts

Showing posts with the label Books and Papers

SureChEMBL: A New Hope

US-D254080-S SureChEMBL has disrupted the field of patent chemistry by liberating chemical structures and knowledge locked in text and images, and by making the compound-patent associations freely  and fully searchable and accessible on a daily basis to everyone: academics, IP professionals, content providers, software vendors, biotechs, small and big pharma, and related chemical industries . The speed, scale and scope of the data is unprecedented for a public resource.  SureChEMBL has been around for less than two years ; during this time, it has evolved into a full-blown chemistry resource provided by the EMBL-EBI: the SureChEMBL interface was revamped and released last year , including combined keyword and structure-based queries against the annotated patent corpus. All chemistry is integrated with UniChem and there are several ways to access the data in bulk, including flat files and a data client. Very soon, the data will be fully integrated and avai...

Paper: Managing expectations: assessment of chemistry databases generated by automated extraction of chemical structures from patents

Our collaborators in GSK have just published an Open Access paper in the Journal of Cheminformatics . It is a comparative study of the quality of chemistry extraction from patent documents and includes patent chemistry sources derived by automated text-mining, such as SureChEMBL and the IBM/NIH data set . Among other things, the paper provides a useful detailed overview of SureChEMBL's chemistry annotation specifications. While conducting this study, we realised that this task is far from trivial for several reasons:  The patent corpus is inherently noisy, ambiguous and error-rich. There are diverse use cases and accuracy expectations when it comes to chemistry extracted from patents. Not all the chemistry found in a patent document is of equal importance. Compound standardisation variants such as stereoisomers, tautomers, salts and mixtures is always an issue. There is a distinct lack of an open Gold Standard when it comes to standardised chemistry extracted fro...

Paper: Activity, assay and target data curation and quality in the ChEMBL database

We've just published an Open Access paper in the Journal of Computer-Aided Molecular Design  on the curation of bioactivity, assay and target data in ChEMBL , including current practices and future plans.  Here is the abstract: The emergence of a number of publicly available bioactivity databases, such as ChEMBL, PubChem BioAssay and BindingDB, has raised awareness about the topics of data curation, quality and integrity. Here we provide an overview and discussion of the current and future approaches to activity, assay and target data curation of the ChEMBL database. This curation process involves several manual and automated steps and aims to: (1) maximise data accessibility and comparability; (2) improve data integrity and flag outliers, ambiguities and potential errors; and (3) add further curated annotations and mappings thus increasing the usefulness and accuracy of the ChEMBL data for all users and modellers in particular. Issues related to activity, assay ...

Paper: PPDMs – A resource for mapping small molecule bioactivities from ChEMBL to Pfam-A protein domains

We've just published a Open Access paper in Bioinformatics on an approach to annotate the region of ligand binding within a target protein. This has a lot of applications in the use of ChEMBL , in particular providing greater accuracy in mapping functional effects, improving ligand-based target prediction approaches, and reducing false positives in sequence/target searching of ChEMBL. Where next for this work - well annotating to a site-specific level would be a good thing to implement (think about HIV-1 RT with the distinct nucleoside and non-nucleoside sites). Here's the abstract... Summary : PPDMs is a resource that maps small molecule bioactivities to protein domains from the Pfam-A collection of protein families. Small molecule bioactivities mapped to protein domains add important precision to approaches that use protein sequence searches alignments to assist applications in computational drug discovery and systems and chemical biology. We have previously propos...

Papers: Literature text mining and extensions to UniChem

Two new papers from the group have just been published, both in Journal of Chemoinformatics - and of course both Open Access. The first deals with some extensions to UniChem to allow far more flexible searches. The abstract is: UniChem is a low-maintenance, fast and freely available compound identifier mapping service, recently made available on the Internet. Until now, the criterion of molecular equivalence within UniChem has been on the basis of complete identity between Standard InChIs. However, a limitation of this approach is that stereoisomers, isotopes and salts of otherwise identical molecules are not considered as related. Here, we describe how we have exploited the layered structural representation of the Standard InChI to create new functionality within UniChem that integrates these related molecular forms. The service, called ‘Connectivity Search’ allows molecules to be first matched on the basis of complete identity between the connectivity layer of their co...

Citing ChEMBL, and Data DOIs

There are now multiple formats and ways to access the ChEMBL data, and we have recently assigned DOIs to all available versions of ChEMBL (and will archive these on the ftp server, permanently). So when you publish use of ChEMBL, could you reference the following papers: ChEMBL Database A. Gaulton, L. Bellis, J. Chambers, M. Davies, A. Hersey, Y. Light, S. McGlinchey, R. Akhtar, A.P. Bento, B. Al-Lazikani, D. Michalovich, & J.P. Overington (2012) ‘ChEMBL: A Large-scale Bioactivity Database For Chemical Biology and Drug Discovery’ Nucleic Acids Res. Database Issue , 40 D1100-1107. DOI:10.1093/nar/gkr777 PMID:21948594 A.P. Bento, A. Gaulton, A. Hersey, L.J. Bellis, J. Chambers, M. Davies, F.A. Krüger, Y. Light, L. Mak, S. McGlinchey, M. Nowotka, G. Papadatos, R. Santos & J.P. Overington (2014) ‘The ChEMBL bioactivity database: an update’ Nucleic Acids Res . Database Issue , 42 1083-1090. DOI:10.1093/nar/gkt103 PMID: 24214965 myChEMBL R. Ochoa, M. Davies...

Paper: An atlas of genetic influences on human blood metabolites

We got involved in the analysis of a really interesting GWAS/metabolomics study, with a publication just appearing in Nature Genetics. A link to the paper is here . Genome-wide association scans with high-throughput metabolic profiling provide unprecedented insights into how genetic variation influences metabolism and complex disease. Here we report the most comprehensive exploration of genetic loci influencing human metabolism thus far, comprising 7,824 adult individuals from 2 European population studies. We report genome-wide significant associations at 145 metabolic loci and their biochemical connectivity with more than 400 metabolites in human blood. We extensively characterize the resulting in vivo blueprint of metabolism in human blood by integrating it with information on gene expression, heritability and overlap with known loci for complex disorders, inborn errors of metabolism and pharmacological targets. We further developed a database and web-based resources for data...

Paper: Towards predictive resistance models for agrochemicals by combining chemical and protein similarity via proteochemometric modelling

There's an Open Access opinion paper out in J. Chem. Biol. on the potential application of chemo/bio joint QSAR techniques that have historically been developed and applied to primarily human pharmaceutical applications, to the agrochemical area. Link to the paper is here . Later in the summer of 2014, there is some very exciting news on agrochemical data for ChEMBL! %0 Journal Article %D 2014 %@ 1864-6158 %J Journal of Chemical Biology %R 10.1007/s12154-014-0112-2 %T Towards predictive resistance models for agrochemicals by combining chemical and protein similarity via proteochemometric modelling %U http://dx.doi.org/10.1007/s12154-014-0112-2 %I Springer Berlin Heidelberg %8 2014-05-15 %K Polypharmacology %K Cheminformatics %K Machine learning %K Resistance %A Westen, Gerard J.P. %A Bender, Andreas %A Overington, John P. %P 1-5 %G English

Paper: Chemical, Target, and Bioactive Properties of Allosteric Modulation

We have just had a paper accepted in PLoS Computational Biology on the work we've done on allosteric modulators (first mentioned on the blog  here ).  The work is based on the mining of allosteric bioactivity points from ChEMBL_14. The data set of allosteric and non-allosteric interactions is available on our FTP site ( here ). This blogpost will just highlight some sections of the paper, but we would like to refer the interested reader to the full paper ( here ).  Dataset The dataset contains ChEMBL annotated and cleaned data divided in both an 'allosteric' set and a 'non-allosteric' (or background) set. Abstracts and titles mentioning allosteric keywords were pulled and from the resulting papers we extracted the primary target and all bioactivities on this primary target. From the remainder of the papers we also retrieved the primary target and all bioactivities on this primary target in a similar manner.  Targets When we observed the target distr...

Paper: The ChEMBL database: a taster for medicinal chemists

We have just had this editorial published in Future Medicinal Chemistry . It is a high-level overview of ChEMBL , SureChEMBL and related resources, aimed primarily at medicinal chemists. The Open Access paper is here . %T The ChEMBL database: a taster for medicinal chemists %J Future Medicinal Chemistry %D 2014 %O DOI: 10.4155/fmc.14.8 %A G. Papadatos %A J.P. Overington

Paper: myChEMBL - a virtual machine implementation of open data and cheminformatic tools

We have just had a paper published in Bioinformatics on myChEMBL - the Linux VM that contains a fully functional version of the ChEMBL database. The paper is here . myChEMBL is available for download at: ftp://ftp.ebi.ac.uk/pub/databases/chembl/VM/myChEMBL/current A warning, it is a fairly big download ( ca. 18GB, so try and do this over a fast stable connection) Source code is available here: https://github.com/rochoa85/myChEMBL %T myChEMBL: A virtual machine implementation of open data andcheminformatics tools %J Bioinformatics %D 2013 %O DOI:10.1093/bioinformatics/btt666 %A M. Davies %A G. Papadatos %A F. Atkinson %A J.P. Overington

Paper: The ChEMBL bioactivity database: an update

An update to what has happen to the Wellcome Trust funded database  ChEMBL  over the past few years has just been published - it seems odd, that we've been around long enough to achieve our 2nd NAR Database paper - so much more to do though! This paper contains features and content up to ChEMBL 17. This could put you in a difficult position which NAR paper to cite in your own publications using ChEMBL; so we suggest both! ;) Oh, and it's Open Access, of course. %J Nucleic Acids Research %D 2013 %P 1–8 %O doi:10.1093/nar/gkt1031 %T The ChEMBL bioactivity database: an update %A A.P. Bento %A A. Gaulton %A Anne Hersey %A L.J. Bellis, %A J. Chambers %A M. Davies %A F.A. Krueger %A Y. Light %A L. Mak %A S. McGlinchey %A M. Nowotka %A G. Papadatos %A R. Santos %A J.P. Overington jpo

Paper: The Functional Therapeutic Chemical Classification System

Here 's an Open Access paper from Samuel in the group. Drug repositioning is the discovery of new indications for compounds that have already been approved and used in a clinical setting. Recently, some computational approaches have been suggested to unveil new opportunities in a systematic fashion, by taking into consideration gene expression signatures or chemical features for instance. We present here a novel method based on knowledge integration using semantic technologies, to capture the functional role of approved chemical compounds. In order to computationally generate repositioning hypotheses, we used the Web Ontology Language (OWL) to formally define the semantics of over 20,000 terms with axioms to correctly denote various modes of action (MoA). Based on an integration of public data, we have automatically assigned over a thousand of approved drugs into these MoA categories. The resulting new research resource is called the Functional Therapeutic Chemical C...

Paper: Target Prediction for an Open Access Set of Compounds Active against Mycobacterium tuberculosis

Here's a paper detailing some multi-method target prediction work as part of the GeMoA FP7 project. Proud, as ever, to publish Open Access. %A Martínez-Jiménez F %A Papadatos G %A Yang L %A Wallace IM %A Kumar V %A Pieper U %A Sali A %A Brown JR %A Overington JP %A Marti-Renom MA %D 2013 %T Target Prediction for an Open Access Set of Compounds Active against Mycobacterium tuberculosis %J PLoS. Comput. Biol. %V 9 %P e1003253 %O doi:10.1371/journal.pcbi.1003253 jpo

Congratulations to Ben Stauch PhD!

Ben Stauch in the group has just been examined on his these - ' Methods for the Investigation of Protein-Ligand Complexes '. This was a tour de force of many techniques - NMR, computational and X-ray crystallography. Ben will be around for a few more months, writing things up, and completing/starting some experimental work on Xe complex refinement and characterisation. Congratulations to Ben from all the group! In due course, the thesis will be downloadable from the EBI and EMBL websites, and I'll update this post when the files are there. jpo

Document Similarity in ChEMBL - 2

Following up on yesterdays post by George and Mark, I put together a slide, hopefully illustrating the advantages of document comparison using objects other than words alone. jpo

Paper: Benchmarking of protein descriptor sets in proteochemometric modeling (part 2): modeling performance of 13 amino acid descriptor sets

A paper from Gerard in the group on some of his proteochemometric modelling work; a link to the paper is here . Z-scales rule! (the original Sandberg et al   J Med Chem paper on the Z-scales was one of my 'lightbulb turning on' moments in my professional life - go hunt it down if you don't know it.) %T Benchmarking of protein descriptor sets in proteochemometric modeling (part 2: modeling performance of 13 amino acid descriptor sets %A G.J.P. van Westen %A R.F. Swier %A I. Cortes-Ciriano %A J.K. Wegner %A J.P Overington %A A.P. IJzerman %A H.W.T. van Vlijmen %A A. Bender %J J. Cheminformatics %D 2013 %V 5 %O doi:10.1186/1758-2946-5-42 jpo

What is the R&D Cost of a New Medicine?

Here 's a recent (2012), and excellent, analysis and estimate of the development costs of a new medicine (specifically an NME, a chemically distinct, novel molecule). There is a good overview of the historical trends in costs and attrition, and a collection of all significant previous estimates of the R&D costs of a new drug. There's some nice exploration of the sensitivity of the costs to various factors, and differential success and costs across various therapeutic areas. In case you wanted to jump to the punchline, the costs in this study is $1,506,000,000 (i.e. $1.5bn) at 2011 USD prices. The report is free, with only registration at the OHE website required to download the report. Great value! %T The R&D Cost of a New Medicine %A J. Mestre-Ferrandiz %A J. Sussex %A A. Towse %I Office of Health Economics %D 2012 %O ISBN 978-1-899040-19-3 jpo

Paper: The Open Access Malaria Box: A Drug Discovery Catalyst for Neglected Diseases

Historically, one of the key problems in neglected disease drug discovery has been identifying new and interesting chemotypes. Phenotypic screening of the malaria parasite, Plasmodium falciparum has yielded almost 30,000 submicromolar hits in recent years. To make this collection more accessible, a collection of 400 chemotypes has been assembled, termed the Malaria Box. Half of these compounds were selected based on their drug-like properties and the others as molecular probes. These can now be requested as a pharmacological test set by malaria biologists, but importantly by groups working on related parasites, as part of a program to make both data and compounds readily available. In this paper, the analysis and selection methodology and characteristics of the compounds are described. Data from studies involving the Malaria Box will be in ChEMBL and the ChEMBL Malaria portal. %A T. Spangenberg %A J.N. Burrows %A P. Kowalczyk %A S. McDonald %A TNC Wells %D 2013 %T The Open Ac...