Showing posts with label cheminformatics. Show all posts
Showing posts with label cheminformatics. Show all posts

Wednesday, February 17, 2010

While we predict drug targets, pharma already knows them

All the excitement about predicting drug targets from remote chemical structure similarities (Nature) and drug side effects (Science) suddenly seems strangely unfounded when you realize that "they" already have the answer:
This figure is a tiny section of the data used for Preclinical Safety Profiling at the Novartis Institutes for BioMedical Research. They are not alone: BioPrint from Cerep is a repository of 2500 compounds assayed against 159 targets. Paolini et al. mention 600,000 binding data points at Pfizer.

Things are begining to change, thanks to academic curation efforts like BindingDB, PDSP Ki database, DrugBank etc. and not least to EMBL-EBI's acquisition of ChEMBL. Still, one can only imaging how different current academic research questions would be if the all-against-all drug–target assay information would be public. In essence, academia is working on predicting additional data points from a sparse matrix, while pharma companies already have the full matrix in hand, but want to add more rows/columns to the matrix in silico.

Update: As always, good discussion on FriendFeed.

Monday, November 10, 2008

4th German Conference on Chemoinformatics: ChEMBL

The talks of the first full day of the 4th German Conference on Chemoinformatics are over. Most interesting for me was Christoph Steinbeck's talk about the recently announced data acquired by the EBI. The database will be called "ChEMBL". There will be a monthly update cycle, so the acquisition does not only capture the current state, but the database is going to be extended. There are three parts (although they'll be combined eventually):
  • "DrugStore": interactions for 1500 drugs. Christoph says that he doesn't expect this to go much beyond what's already publicly available in DrugBank et al. today.
  • "CandiStore": 15,000 clinical leads
  • "StARLite": 500,000 medical chemistry leads. This is where most of the novelty (in terms of public data) lies. For this part, there are >5500 annotated targets, >3500 of which are proteins (the rest is e.g. tissues), and 2 million experimental bioactivities. The database contains bidirectional links to the literature on synthetic routes and assays for the ligands and descriptions of the targets.
The data will be first made available as database dumps, more user-friendly interfaces will be added later.

Two URLs of interest that I didn't know before: The ChEMBL blog and John Overington's lab homepage.

Other remarks about today will follow when I have a real internet connection (not just 6 kB/s via Bluetooth/GPRS for 9 ct/min) to do some more background research.