search

Tampilkan postingan dengan label web service. Tampilkan semua postingan
Tampilkan postingan dengan label web service. Tampilkan semua postingan

Jumat, 15 Oktober 2010

Dynamic links to private tagged Mendeley collections

My close collaborators and I have been using Mendeley as a convenient way to share PDFs of journal articles. Not all of us have access to the same libraries so links are not enough - we need the full documents. We also use Dropbox as a redundancy but Mendeley allows tagging and recording notes, which is very handy for everyone in the group.

Now that Mendeley is providing an API, Andrew Lang has written code that significantly leverages the information in our private ONS collection. We can now create public links that return the most updated results for specific tags, including multiple tags (which I don't think you can do on Mendeley). For example the following link returns all articles in the ONS collection tagged with "science2.0" and "chemistry":
The results include available information from Mendeley, including the title, authors, journal citation, doi, url, tags and the abstract. Because this information is public the PDFs can't be provided but the hyperlinks make it as convenient as possible.


At the end of the report the full list of all available tags for the ONS collection is provided. A more refined or different search can be done immediately simply by checking boxes and hitting the submit button.Because the tags are controlled by the users of the private collection, these links can be useful when discussing an ongoing project and referring to a very specific topic. For example, we have been collecting examples of articles where a Ugi reaction is carried out and the product precipitates. This link provides an updated report on that very narrow topic:
http://showme.physics.drexel.edu/onsc/mendeley/?tags=Ugi+precipitate
There are still 2 major limitations to this service:

1) The search is very slow (can take a minute or two) because there is no way currently to use the Mendeley API to selectively return results based on tags. Every search requires initially returning all results for the collection (currently a few hundred).

2) Notes are currently not returned. If the API is updated to include these the usefulness would increase dramatically. For example in the results for the above query I took notes of the conditions involved in the Ugi precipitate for each paper. With the current format, one has to read each paper to find the relevant information.

Progress on our Mendeley related services will be posted on the ONSwebservices wiki.

Senin, 23 Agustus 2010

ChemTaverna Workflows of ONS Web Services now on MyExperiment

I'm pleased to report that one of the collaborations initiated at the Berkeley Open Science conference last month is progressing very well.

Carole Goble introduced me to Peter Li who runs the ChemTaverna project. The idea was to use Taverna to construct workflows using the web services developed by Andrew Lang for our Open Notebook Science projects: UsefulChem and the ONS Solubility Challenge.

Peter quickly created several workflows to demonstrate what is possible. Here is a workflow that uses a Google Spreadsheet as input. SMILES for amines, carboxylic acids, aldehydes and isonitriles are entered in the appropriate columns. The workflow first creates a virtual library of Ugi products from all possible combinations of reactants. Then each product is submitted to a web service that predicts the solubility in methanol, the most common solvent for Ugi reactions.

The resulting spreadsheet can then be sorted by predicted solubility to recommend products that are more likely to precipitate from the reaction mixture. In this particular example Ugi products derived from boc-glycine are predicted to have a low solubility in methanol. The least soluble compound is predicted to have a solubility of only 0.07MIn this library, Ugi products derived from boc-methionine are predicted to be too soluble to precipitate. For example this Ugi product has a predicted solubility of 3.7 M.
(note: ChemSpider has a tendency to draw the minor tautomer for some amides and carbamates)

There are a few issues to take into consideration in order to use this particular workflow:

1) This will only work on Taverna Workbench 2.1.2 with these plug-ins installed. At one point it will be made to work on Taverna Workbench 2.2 and uploaded onto MyExperiment. The workflow used here is currently available here.

2) The SMILES in the input Google Spreadsheet must be written in the format of the current example (aldehyde, amine and isonitrile groups on the left and carboxylic acid groups on the right)

3) All of the Ugi products in the virtual library must already exist in ChemSpider. Otherwise, the solubility predictions will fail because of missing descriptors as discussed previously.

Peter has uploaded simpler workflows onto MyExperiment that are compatible with the current version of Taverna Workbench (v2.2).

First, the generation of Ugi product libraries from reactant SMILES in a Google Spreadsheet is available here.

Another workflow handles the prediction of Abraham descriptors.
This workflow processes the prediction of solubility for a given solute and solvent.

The main rationale for incorporating web services derived from our Open Notebook Science projects into Taverna is leverage. MyExperiment already benefits from a vigorous community of developers in the bioinformatics arena. With the growth of the ChemTaverna initiative, the integration of cheminformatics and bioinformatics workflows should become seamless.

By making our solubility and chemical reaction web services available in formats that are convenient for others to use it increases the opportunities that our work will be actually useful. It also makes it easier for us to leverage the resources made available by others for our own applications in drug discovery and reaction design.

Essentially this means that we have extended the reach of the information cascade triggered by the recording of an experiment in a laboratory notebook and a very simple abstraction process to represent that experiment in a semantically addressable format.

Minggu, 25 Juli 2010

General Transparent Solubility Prediction using Abraham Descriptors

Making solubility estimations for most organic compounds in a wide range of solvents freely available has always been a main long term objective for the Open Notebook Science Solubility Challenge. With current expertise and technology, it should be as easy to obtain a solubility estimate as it is now to get driving directions off the web.

Obviously this won't be attained purely by exhaustive measurements, although we have been focused on strategic measurements over the past two years. In parallel, we have been constantly evaluating the various solubility models out there for suitability.

Although there are several solubility models available for non-aqueous solvents, our additional requirement for transparent model building has proved surprisingly difficult to satisfy.

From this search, the Abraham solubility model [Abraham2009] floated to the top, with an important factor being that Abraham has made available extensive compilations of descriptors for solutes and solvents. In addition the algorithms used to convert solubility measurements to Abraham descriptors (a minimum of 5 different solvents per solute) has allowed us to generate our own Abraham descriptors automatically simply by recording new measurements into our SolSum Google Spreadsheet. These can be obtained in real time as well.

This approach permitted us to provide predictions for a limited number of solutes in a wide range of solvents and we have included these predictions in the past two editions (2nd and 3rd) of the ONS Challenge Solubility Book.

Coming at the problem from a different approach, Andrew Lang has also been trying to predict solubility using only open molecular descriptors, mainly relying on the CDK. Since our most commonly used solvent has been methanol, Andy recently generated a web service to predict solubility in that solvent.

By combining these two approaches, Andy has now created a modeling system that can not only generally predict solubility in a wide range (70+) of solvents - but it can also provide related data that can be used for modeling other phenomena such as intestinal absorption of a drug or crossing the blood-brain barrier.[Stovall 2007]

The idea is to use a Random Forest approach to select freely available descriptors to predict the Abraham descriptors for any solute. A separate service then generates predicted solubilities for a wide range of solvents based on these Abraham descriptors. I'm using the term "freely available" because - although the CDK descriptors and VCCLab services are open - the model requires 2 descriptors only available from ChemSpider (ultimately from ACD/Labs).

Here is an example with benzoic acid. As long as the common name resolves to a single entry on ChemSpider, it is enough to enter it and it automatically populates the rest of the fields, which are then used by the service to generate the Abraham descriptors.

Hitting the prediction link above will automatically populate the second service and generate predicted solubilities for over 70 solvents.

This approach of allowing people to access these components separately can be useful. It can be instructive to manually play with the Abraham descriptors directly to see how predicted solubilities are affected. There are also situations where one has experimentally determined Abraham descriptors for a solute and bypassing the descriptor prediction step is required.

However, for those who prefer to cut to the chase, a convenient web service is available where the common name (or SMILES) of the solute is entered and the list of available solvents appears as a drop down menu.

Now here is where I think the real payoff comes for accelerating science with openness. Andy has also created a web service that returns the predicted solubility in molar as a number from common names (or SMILES) for solute and solvent via the URL. For example click this for benzoic acid in methanol. The advantage here is that solubility prediction can be easily integrated as a web service call from intuitive interfaces such as a Google Spreadsheet to enable even non-programmers to make use of the data. Notice that the web service provided in the fourth column for the average of measured solubility values enables an easy way to explore the accuracy of specific predictions.

Such web services could also be integrated with data from ChemSpider or custom systems. If those who use these services feed back their processed data to the open web, it could take us a step closer to automated reaction design. For example consider the custom application to select solvents for the Ugi reaction. Model builders could also use the web services for predicted and measured solubility directly.

A while back we explored using Taverna for MyExperiment to create virtual libraries of SMILES. Unfortunately we ran into issues with getting the applications developed on Macs to run on our PCs. This might be worth revisiting as a means of filtering virtual libraries through different thresholds of predicted solubility.

Andy has described his model in detail in a fully transparent way - the model itself, how it was generated and the entire dataset can be found here. We would welcome improvements of the model as well as completely new models based on our dataset using only freely available tools.

It should be noted that when I use term "general" it refers to the ability for the model to generate a number for most compounds listed in ChemSpider. Obviously compounds that most closely resemble the training set are more likely to generate better estimates. Because of our synthetic objectives using the Ugi reaction we have mainly focused on collecting solubility data for carboxylic acids, aldehydes and amides either from new measurements or from the literature.

Another important point concerns the main intended application of the model: organic synthesis. Generally the range of interest for such applications is about 0.01 - 3M. This might be very different for other applications - such as the aqueous solubility of a drug, where distinctions between much lower solubilities may be important.

For a typical organic synthesis, a solubility of 0.001M or 0.005M will probably translate as effectively insoluble. This might be a desired property for a product intended to be isolated by filtration. On the other end of the scale knowing that a solubility is either 4M or 6M will not usually have an impact on reaction design. It is enough to know that a reactant will have good solubility in a particular solvent.

Given the above considerations for intended applications and the likelihood that the current model is far from optimized, the predictions should be used cautiously. We suggest that the model is best used as a "flagging device". For example, if a reaction is to be carried out at 0.5M, one may place a threshold at 0.4M for the predicted values of reactants during solvent selection, with the recognition that a predicted 0.4M may be an actual 0.55M. A similar threshold approach can be used for the product, where in this case the lowest solubility is desired. A practical example of this is the shortlisting of solvents candidates for the Ugi reaction.

Another example of flagging involves identifying the outliers in the model. These can be inspected for experimental errors and possibly remeasured. Alternatively outliers may shed light on the limitations of the model. For example we have found that the solubility of solutes with melting points near room temperature can be greatly underestimated by the current model. This may be an opportunity to develop other models which incorporate melting point or enthalpy of fusion.[Rohani 2008]

Although it is possible that better models and more data will improve the accuracy of the predictions, this can be true only if the training set is accurate enough. Based on conversations I've had with researchers who deal with solubility, reading modeling papers and our own experience with the ONS Challenge I am starting to suspect that much of the available data just isn't accurate enough for high precision modeling. Models using data from the literature are especially vulnerable I think. Take a look at this unsettling comparison between new measurements and literature values (not to mention the model) for common compounds.[Loftsson 2006] Here is a subset:
I have also made the point in detail for the aqueous solubility of EGCG. Could this be the reason that so many different solubility models using different physical chemistry principles have evolved and continue to co-exist?

The situation reminds me a lot of the discussions taking place in the molecular docking community.[Bissantz 2010] The differences in calculated binding energies are often small in comparison with the uncertainties involved. But docking can still be used as one tool among others to find drug candidates by flagging a collection of compounds above a certain threshold binding energy.

Jumat, 30 April 2010

NMR integration web service expanded

The ONS Challenge has extensively used a web service created by Andrew Lang to automatically calculate solubility from NMR spectra. One of the constraints of the service was that the JCAMP-DX file had to be deposited in a special folder on a server at Drexel.

Andy has now modified the script so that the JCAMP-DX file can be located anywhere on the internet. I have prepared a modified Google Spreadsheet to serve as a template for SAMS calculations (Semi-Automated Measurement of Solubility). Simply enter the url to the JCAMP-DX file in the appropriate column and fill in the ppm ranges and corresponding hydrogen numbers for the solvent and solute, and molecular weight and density data. (The predicted density of solids can be found on Chemspider). The concentration of the solute will then be automatically calculated based on an assumption of volume additivity.

The web service (which handles baseline correction) could be used for any other purpose involving the integration of spectra. Just make a copy of the Google Spreadsheet and modify.

Note that the JCAMP-DX files must be in XY format. If your instrument saves spectra in a compressed format they must be converted to XY. The desktop version of Robert Lancashire's JSpecView can be used to carry out the conversion.

This template spreadsheet also features a service in a cell to display the NMR spectrum by simply clicking on the link inside the cell. This is very handy because it obviates the need to create an HTML file which must normally accompany the JCAMP-DX file for viewing. Being able to quickly view a spectrum from a particular row within the Google Spreadsheet makes tracking data provenance very intuitive and errors easy to spot.

Rabu, 10 Maret 2010

Updated Chemistry Web Services - now with Density

I mentioned a while back the web services that Rajarshi Guha had set up for us. We are often in need of molecular weight and density data for both solutes and solvents since we rely on an assumption of volume additivity when calculating concentration.

Since Rajarshi moved to the NIH, the location of the services has changed. We now have the CDK installed on a Drexel server so some of the simple services like MW and SMILES generation are still available there.

However density has been challenging to provide as a service. Experimental density values for solvents are commonly available but the calculated densities of solids is hard to find. ChemSpider is one of the few sources where calculated densities of solids and liquids are freely available. Unfortunately there are currently no ChemSpider density web services.

As an interim solution for the UsefulChem and ONSChallenge projects we have set a look-up table as a Google Spreadsheet (SolventLookUp) for most solvents of potential interest. Solutes added to our SolubilitiesSum sheet are automatically added to a SoluteLookUp SQL database running at Oral Robert University and the ChemSpider densities are added there via an automated but slow process.

Andrew Lang has used these resources to provide web services returning densities and other properties or descriptors. These data sources are especially important for the nearly automated production of new editions of the ONS Challenge Solubility Book. This is not a general solution since it only includes compounds of interest to our group and would not scale (at least for licensing reasons) to millions of compounds.

But it does come in handy for us because we can quickly call these services within a Google Spreadsheet to do a variety of useful calculations, minimizing the possibility of error by copy and pasting.

As an example see the following ChemServices sheet. Enter the common name for a solvent or solute and the number of millimoles and the sheet will automatically calculate the corresponding number of milligrams or microliters. [Note that Google Spreadsheets can only handle a maximum of 50 web service calls at a time - a useful trick is to highlight cells after the calculations then copy and "paste as values". Make sure to keep some cells with the web service calls in case you need to do more calculations in the future]

Senin, 04 Mei 2009

Streamlining automated solubility measurements with NMR JCAMP-DX files

Two months ago I reported on a protocol for measuring solubility using NMR JCAMP-DX files and a web service set up by Andrew Lang called from within a Google Spreadsheet. Things were going well until David at ORU was a little too productive and crashed the server from too many requests.

Andy had to change the way the script worked and used this as an opportunity to make the service more broadly usable. It turns out that the compressed JCAMP-DX files produced by different NMR instruments are not created with exactly the same standards. A way to address that issue is to convert the files to an uncompressed XY format. Unfortunately, there was a glitch in JSpecView which created XY formatted spectra displaying in Hz instead of the standard ppm.

Now all of these issues have been resolved and the process is simpler than ever. Robert Lancashire fixed the glitch in the April 26, 2009 release of JSpecView. And Andy not only made his integration web service work for the new release but also created another service to display JCAMP-DX spectra directly from the the DX file (see here for an example). In the past students had to create an associated HTML file to display JCAMP-DX files and this was just another point in the process to introduce errors and slow things down.

The new process for the semi-automated measurement of solubility (SAMS) using NMR is as follows:

1) Make a saturated solution in a given solvent (sonicate for at least 30 mins - more on this in a separate post)
2) Transfer about 0.1 mL to an NMR tube with some compatible deuterated solvent (for locking)
3) Take the NMR spectrum and export the JCAMP-DX file (on our machines these are in a compressed format)
4) Open the initial JCAMP-DX files in JSpecView and save as JCAMP-DX XY format
5) Upload the converted file to the ONSC server in the spectra folder
6) Fill out the requested information in the SAMS spreadsheet and you have the solubility calculation (first open the SAMStemplate and save as a copy with a new name)

We can now easily finish processing the backlog of measurements that the ONSchallenge participants have been obtaining and record them in the SolubilitySum spreadsheet for querying.