Skip to main content

Posts

Showing posts with the label ChEMBL Curation Notes

Restructuring of ACTIVITY data: VALUE, TEXT_VALUE and ACTIVITY_COMMENT

ChEMBL is a bioactivity database for drug-like compounds. The ACTIVITIES table stores the readout from bioassays that test compounds against a biological target. These are often numerical readouts such as Ki, IC50, Inhibition %, or Cmax but are occasionally non-numerical summary data, e.g., “Not Soluble”, “Not active”, or “Active”. Historically, the non-numerical data was captured as an ACTIVITY_COMMENT but since ChEMBL 24 this has been more accurately captured as TEXT_VALUE.  Whilst the TEXT_VALUE field has been used more extensively in recent releases, legacy data covering experimental outcomes, observations, context such as threshold values, and other metadata is still largely hosted in the ACTIVITY_COMMENT field. Moving forward, only the TEXT_VALUE field will be used to report the primary outcome of an experiment where the output is not numerical, for example categorical data.  This could be reporting an Activity (e.g., “Active”/ “ Not active”), Toxicity (e.g., “Toxic”/ “ ...

Streamlining the Pesticide Mechanism of Action Classification

Since ChEMBL 20, molecules in ChEMBL that are known pesticides have been linked to the mechanism of action classification assigned by the Fungicide Resistance Action Committee (FRAC: http://www.frac.info), Herbicide Resistance Action Committee (HRAC: http://www.hracglobal.com) or Insecticide Resistance Action Committee (IRAC: http://www.irac-online.org). These Committees provide information on the mechanism of action of key pesticides as part of ongoing efforts to combat pesticide resistant fungi, plants, and insects, and to prolong the use of existing pesticides. Supplementing ChEMBL with curated pesticide data from key organisations complements the wealth of compounds within ChEMBL explored for applications in human health, supporting a move towards a OneHealth approach. These classification schemes group pesticides both by their mode of action and chemical class. From ChEMBL 20 to ChEMBL 35 classifications were stored within ChEMBL in three tables, with three associated mapping tabl...

Paper: Activity, assay and target data curation and quality in the ChEMBL database

We've just published an Open Access paper in the Journal of Computer-Aided Molecular Design  on the curation of bioactivity, assay and target data in ChEMBL , including current practices and future plans.  Here is the abstract: The emergence of a number of publicly available bioactivity databases, such as ChEMBL, PubChem BioAssay and BindingDB, has raised awareness about the topics of data curation, quality and integrity. Here we provide an overview and discussion of the current and future approaches to activity, assay and target data curation of the ChEMBL database. This curation process involves several manual and automated steps and aims to: (1) maximise data accessibility and comparability; (2) improve data integrity and flag outliers, ambiguities and potential errors; and (3) add further curated annotations and mappings thus increasing the usefulness and accuracy of the ChEMBL data for all users and modellers in particular. Issues related to activity, assay ...

Paper: PPDMs – A resource for mapping small molecule bioactivities from ChEMBL to Pfam-A protein domains

We've just published a Open Access paper in Bioinformatics on an approach to annotate the region of ligand binding within a target protein. This has a lot of applications in the use of ChEMBL , in particular providing greater accuracy in mapping functional effects, improving ligand-based target prediction approaches, and reducing false positives in sequence/target searching of ChEMBL. Where next for this work - well annotating to a site-specific level would be a good thing to implement (think about HIV-1 RT with the distinct nucleoside and non-nucleoside sites). Here's the abstract... Summary : PPDMs is a resource that maps small molecule bioactivities to protein domains from the Pfam-A collection of protein families. Small molecule bioactivities mapped to protein domains add important precision to approaches that use protein sequence searches alignments to assist applications in computational drug discovery and systems and chemical biology. We have previously propos...

Compound Curation - The story so far...

As chemical curator for ChEMBL, I spend a lot of time processing, checking and standardising the compounds in the database. I use various pieces of software for this, but mostly it’s  Pipeline Pilot . For those of you who don’t know, Pipeline Pilot is Accelrys’s graphical scientific workflow authoring application, that allows passing hundreds of thousands of compounds through various components to make sure they meet our standards to be loaded into ChEMBL. However, it’s always incredibly useful to utilise other available software in a complementary manner to see if anything may have been missed, could be done in a different way or just to see what alternative results you can get. One such open source software package is Indigo , created by GGA Software Services . On of the web application developers was passing all of the ChEMBL compounds through the standard Indigo loader, via a Python script, during the course of his work, and found that there were about 9,000 compoun...

Removal of Metal-Containing Compounds

Further to my post a few months ago ( To Remove or Not to Remove ) about removing certain problem metal-containing compounds, we have now come up with a plan of what to do. Instead of labeling this curation as ‘removal of inorganics’, or ‘removal of organometallics’, we simply want this to be known as ‘removal of some metal-containing compounds’. The criterion that we used was to exclude a large proportion of compounds that contained a metal, apart from cases where a metal was commonly found as part of a pharmaceutical preparation (e.g. Ranitidine Bismuth Citrate CHEMBL2111286 , Silver Sulfadiazine CHEMBL1382627 , Bacitracin Zinc CHEMBL2096639 ). The reasoning behind the removal of such compounds was that most of these metals are bonded to the rest of the compound components via coordinate bonds. However, due to InChI limitations , there is no way of creating a Standard InChI that retains coordinate bond information. As we use Standard InChI as the ...

To Remove Or Not To Remove - That Is The Question

During the course of standard compound curation, I come across problem inorganic compounds. An example of these are Cisplatin and Transplatin . These compounds only differ in the orientation of their complex bonds but complex bonds cannot be drawn in a standard molfile without causing InChI issues. At the  moment, they are kept separate by showing standard bonds between the Pt, Cl and NH3 in Cisplatin, but we have removed the bonds altogether for Transplatin. This is not an ideal situation, nor an accurate structural representation. Another example is the compound , below left, and how it should look as a complex, right, from the paper: At the moment, there are approximately 1,800 cases like this, which only accounts for 0.15% of the entire ChEMBL compound set. What we are proposing to do is to remove the structures for these complex compounds and to keep only their names and all of the associated biological data. This would then treat them in a similar way to the a...

ChEMBL Compound Clean Up

For the last three months, I've been busy working my way through a 9000 long (sometimes headache-inducing) set of ChEMBL compound ids. These had been highlighted for curation for the reason that for each ChEMBL_id in the list, there were two or more compound keys from the same paper. This implied that either there were two indistinguishable using InChI representation compounds described in the paper or they were different compounds that had been somehow merged together in the database. Each ChEMBL_id was individually checked against the data in the original paper to see if there were indeed two compound keys for the same structure. The outcome of this check gave rise to one of four cases: The structure(s) was found to be incorrect and was redrawn. The structure was correct for some records but not others, so a new compound was created for those selected records. The structure required the definition of stereochemistry or a salt. The structure was le...

Latest activities on the Activities table in ChEMBL_15

For the recent ChEMBL_15 release, a considerable part of our efforts was focussed on the standardisation and harmonisation of the data in the Activities table. The latter holds all the quantitative and qualitative experimental measurements across compounds, assays and targets; needless to say that without it there's no ChEMBL ! This is a summary of what we've incorporated so far: Flag missing data: Records with null published values and null activity comments were flagged as missing. Standardise activity types and units: Conversion of heterogenous published activity type descriptions and units to a standard_type and set of standard_units (e.g., for IC50 convert mM/uM/pM measurements to nM). Flag unusual units: Records with unusual published units for their respective activity types were flagged as 'non standard'. For example, a hypothetical record with IC50 type and units in kg would be flagged! Convert the log values: The records with activity types ...

Compound Clean Up and Mapping (Posted by Louisa)

This new blog post has been created due to popular demand and user requests. I hope that this is useful for you. After being manually extracted from the primary literature, a compound can be only loaded into the ChEMBL database after it has been run through our in-house clean-up protocol. This protocol utilises Accelrys 's Pipeline Pilot software and has evolved a lot over the past three years. The clean-up protocol is used to prevent any structures from being loaded that could be incorrect, not properly charged or contain bad valences. We also use it to map the structures to already existing compounds in ChEMBL. Historically, the clean-up protocol was very simple with just a few components to squeeze out any unwanted structures. Initially, we were mostly concerned about having uneven charges (e.g. charged counter ion but neutral parent) or quaternary nitrogen-containing compounds without a counter ion at all. Over the past 14 releases, the clean-up has become more sophisti...

ChEMBL Is Alive! Part 1 - posted by Louisa

'ChEMBL Is Alive' is to show that ChEMBLdb is a living database that is constantly being worked on by a number of people. As the Chemical Curator for ChEMBL, I (Louisa Bellis) thought it would be interesting for our Blog readers to find out what goes on behind the scenes at 'ChEMBL Towers' and to get regular updates on what we are doing to the data between releases and in response to user emails sent to chembl-help@ebi.ac.uk. As well as being the chemical curator, I also deal with most of the help-desk  traffic, where users can email in and let us know of any errors that they may have found, or even to suggest an improvement or enhancement for the interface. As an example of the work that is done to ChEMBL on an ongoing basis, I thought it would be good to give a brief summary of some of the chemical curation that occurred during the month of June 2012: An external user pointed out to me that they had come across a 'few' compounds that had the same ca...