Skip to main content

Posts

Showing posts with the label UniChem

Finding Compounds in Databases using UniChem

Have you ever identified an interesting compound and wondered what else is known about it?   For example is there any bioactivity data on it in ChEMBL or PubChem?   Is there any toxicity data on it (CompTox)?   Then having found interesting data on a compound wondered if it can be purchased or whether it has been patented.   All this can be done using UniChem.   Interested?    Come along to our webinar on 29th March at 2pm BST  (3pm CEST, 9am EDT) You will however need to register by emailing chembl-help . Places are limited so please let us know as soon as possible if you register but are then unable to attend. If you want to know more about UniChem please read on. UniChem ( https://www.ebi.ac.uk/unichem/  is a simple system we have developed to cross-reference compounds across databases both internal to EMBL-EBI and externally. Currently we have cross-references to 140 million compounds in 30 different databases. Information about t...

Papers: Literature text mining and extensions to UniChem

Two new papers from the group have just been published, both in Journal of Chemoinformatics - and of course both Open Access. The first deals with some extensions to UniChem to allow far more flexible searches. The abstract is: UniChem is a low-maintenance, fast and freely available compound identifier mapping service, recently made available on the Internet. Until now, the criterion of molecular equivalence within UniChem has been on the basis of complete identity between Standard InChIs. However, a limitation of this approach is that stereoisomers, isotopes and salts of otherwise identical molecules are not considered as related. Here, we describe how we have exploited the layered structural representation of the Standard InChI to create new functionality within UniChem that integrates these related molecular forms. The service, called ‘Connectivity Search’ allows molecules to be first matched on the basis of complete identity between the connectivity layer of their co...

Should CAS numbers be in ChEMBL and/or UniChem?

A very quick survey to add excitement to either your holiday or work-day! None of these sucker links, where there appears a 0.24% complete progress bar on the second page, it's just a simple yes/no question on whether it's a good idea to add CAS registry numbers to ChEMBL and/or UniChem. No promises that we could deliver this, but depending on what you vote for, we will consider our options. Update : Given the multiple channels out there, there are also comments on this on LinkedIn (in the ChUG - "ChEMBL User Group" group - why not join, if you're not already) and a couple on Google+. Update 2 : I'll let the poll run till the end of the week (Friday 8th 2014) - and then write something up on the results.

UniChem Released

For data managers of chemistry resources, the maintenance of structure-based links to other chemistry resources can be a tedious chore. The job is all the more burdensome knowing that your counterparts in other chemistry based-resources are essentially duplicating your efforts, in order to keep their links to your resource updated. In an attempt to remove this duplication of effort, and automate the processes involved, we have developed UniChem ,  and which is described in a recent publication . Getting structure-based links out of UniChem can be achieved either via the web-interface or the web services. For automated updating, using the web-services is often the best choice. The current set of web service methods has been designed to allow users several options for how they might obtain links data. Below are detailed two possibilities. One such option would be to use the following methods: First, query UniChem for all valid src...

UniChem - An EBI compound structure cross-referencing resource

We have faced for some time some issues with compound integration with ChEMBL - specifically the loading of compound sets into ChEMBL for cross referencing, between for example, ChEBI, PDBe compounds, etc . The ChEMBL update cycle is relatively slow with respect to some other resources, and there is inevitable thrash with compounds not being present, especially for exciting new data. Without doing something different for compound integration, we were starting to face a scenario where we had a compound table with many millions of compounds without any bioactivity data, and following this the inevitable slowdown in searching, etc . We also had some issues facing us about curation of other people's primary data, changing compound structures, or their rendering, etc . So, we decided to set up an external system to resolve cross-references between various databases. This is a very simple Standard InChI lookup, containing compounds from resources such as ChEMBL, ChEBI, PDBe, Dru...