search

Tampilkan postingan dengan label crowdsourcing. Tampilkan semua postingan
Tampilkan postingan dengan label crowdsourcing. Tampilkan semua postingan

Senin, 17 Agustus 2009

My first talk at ACS09 fall meeting on Crowdsourcing Solubility and ONS

Yesterday (August 16, 2009) I gave my first talk at the ACS meeting in Washington. It was part of an outstanding session on Chemical Text Mining and Public Molecular Databases, organized by Antony Williams and Alex Tropsha.
9:00 AM 1 U.S. EPA computational toxicology programs: Central role of chemical-annotation efforts and molecular databases
Ann M. Richard, Maritja A. Wolf, ClarLynda R. Williams-Devane, Richard Judson
9:25 AM 2 Linking public and commercial chemical data: ChemSpider and SureChem
Nicko Goncharoff
9:50 AM 3 Building an integrated system for chemistry markup and online publishing integrated to online chemistry resources
A J Williams
10:30 AM 4 Turning mining inside out
Colin R Batchelor
10:55 AM 5 Chemreader: A tool for extracting chemical structure information from digital raster images
Jungkap Park, Kazu Saitou, Kerby Shedden, Gus R. Rosania
11:20 AM 6 Exploiting a hidden treasure: Automated chemical entity recognition in Chemisches Zentralblatt
Valentina Eigner-Pitto, Heinz Saller, Peter Loew
1:30 PM 12 Online chemical modeling environment: database
Sergii Novotarskyi, Iurii Sushko, Robert Körner, Anil Kumar Pandey, Igor V. Tetko
1:55 PM 13 Public molecular databases: How can their value be increased by generation of additional data in silico?
Vladimir V. Poroikov, Dmitry Filimonov, Marc C. Nicklaus
2:20 PM 14 Chemical space management of large libraries for new active small molecules selection for prostate cancer treatment
Andrew V. Scorenko, Andrei A. Gakh, Andrey V. Sosnov, Mikhail Yu. Krasavin
2:45 PM 15 Crowdsourcing nonaqueous solubility and synthesis using Open Notebook Science
Jean-Claude Bradley, Khalid Mirza, Rajarshi Guha, Andrew Lang, A. Williams
3:25 PM 16 ChemXSeer: A cyberinfrastructure for environmental chemical kinetics
Karl T. Mueller, William J. Brouwer, C. Lee Giles, Prasenjit Mitra, Carl Lagoze
4:15 PM 18 Reliable reactions and stable structures
Jonathan M Goodman
Many of the presentations highlighted the use of ChemSpider or full collaborations (such as the integration with SureChem patent data). The acquisition of ChemSpider by RSC was repeatedly discussed and this seems to have accelerated such collaborative projects. Colin Batchelor from the RSC provided a great talk on their approach of using ontologies to better leverage the power of chemistry publications. [The presentations were judged and Colin won first prize - I won second, which was pretty cool :) and won me a ticket to the CINF lunch on Tuesday]

I also got to meet Gus Rosania in person for the first time. We had met via the blogsphere a while back over our interests in malaria and Open Notebook Science. Gus was there to share his results from ChemReader, a software package he developed to automatically read chemical structures from images.

I started my presentation by detailing the recent events surrounding the report of the oxidation of secondary alcohols using NaH. The timing of this was perfect because it really showed how useful it can be to immediately share the full data of experiments. This is the type of thing that would have been extremely helpful during the initial reports of Cold Fusion but the tools for sharing in such a detailed way were just not available. Carmen Drahl just wrote an article about this for the August 17, 2009 issue of Chemical & Engineering News (subscriber access).



Minggu, 05 Juli 2009

Regression of 5D solubility space and distributed automation

I recently reported on the plotting of a solubility surface in 3D. Marshall Moritz has now extended his measurements of the solubility of 4-nitrobenzaldehyde in 2 more mixed solvent systems (ONSC-EXP114), giving us 4 solvents and temperature. The results are stored in the SolSumMix spreadsheet.

Andrew Lang has performed a quadratic regression analysis of this space and we have pretty good agreement with the experimental data points (see the "predicted solubility" column in the above SolSumMix spreadsheet).

Although we can't easily represent the entire 5D space intuitively, we can take 3D slices of the regression to assess the fit. For example, consider the plot of mol fraction % chloroform vs. acetonitrile keeping other solvents at zero concentration. What we observe is a nice saddle shape similar to the plot we did earlier with the original data points.
Now consider a slice of mol fraction % toluene vs THF keeping the other 2 solvents at zero. For temperatures above about 0 C we observe an expected rise in solubility with temperature. However going below 0 C the curve reverses and solubility is predicted to go up a bit. This is clearly not right and it simply means that we are missing key data points in that area. A quadratic fit will insert parabolic elements giving this inversion. It is very important to understand that these models will probably do fairly well for intrapolating within our experimental range (about -25C to 40C) but will not be very helpful for extrapolating beyond this region.
Since it is difficult to manually inspect every possible slice of this 5D space Andy has created a service that returns recommendations for the most needed points to be measured next to generate a better model. We already have a DoSol spreadsheet that instructs ONS Challenge students as to the most urgent next solubility measurement to make. This additional "bot" (not quite fully automated as of yet but will be soon) integrates nicely with the collection we already have.

We aim to show that such open distributed mechanisms to requests and execute measurements is a viable way to efficiently leverage crowdsourcing to automate parts of the scientific method. If it can be applied to solubility it can be applied to other problems.

Minggu, 14 Juni 2009

Crowdsourcing solubility requests from bots and people

Now that we have a reasonably routine way of measuring non-aqueous solubility and students trained to do the measurements, we can think about adding some structure to the workflow. By using an open Google Spreadsheet to list the next solute and solvent to measure we can easily crowdsource the requests. This is the dosol sheet and students are instructed to check it before planning their experiments for the day.

Requested experiments come from the following sources:

1) OUTLIER BOT: Andrew Lang has created a neat service that reports measurements exceeding a provided standard deviation to mean ratio (plug into the URL) and a Grubbs outlier threshold. Currently, the "bot" does not automatically write to the dosol sheet but eventually I see that it would make sense to complete that integration. On a daily basis I run it and process the flagged entries in the following ways:
a) If I investigate the corresponding lab notebook pages and find an error in the calculations I simply fix it.
b) If I find that one or more of the measurements are obviously in error I mark the entries as DONOTUSE in the SolSum sheet. This is often a low value that was obtained early on in the project before we appreciated how much mixing is required for some solutes.
c) If I cannot determine why there is a discrepancy I place a request on the dosol sheet. (note: solubilities that are very low (<0.1 M) may have relatively high standard deviation to mean ratios because the techniques we use are not precise at those solubilities and remeasuring won't help)

2) Internal Requests: If someone from within our group is planning a Ugi reaction and there is a missing solubility measurement for one of the reagents or Ugi product, this is a convenient way to request the missing information to make a solvent choice for the reaction.

3) External Requests: When people think about crowdsourcing, usually the first thing that comes to mind is how others can help solve their problem. But in addition to asking what the science world can do for you why not ask what you can do for the science community? We have done so last week and received 3 requests (via a form created by Andy). One for the solubility of iodine in 1,1,2-trichloroethane, one for the solubility of nitric acid in water and another for pyrene in acetonitrile.

For the first request, iodine has no hydrogens so we can't use our standard NMR technique to do it so I set the priority of that request lower. For nitric acid - it is well known that it is miscible in water so we won't be doing that experiment. We also want to focus on non-aqueous solubilities. However the third request was a good fit and Marshall Moritz from our lab processed it right away - he found a value on 0.08M for pyrene in acetonitrile at room temperature (EXP108).

Hadas Joseph from the Agricultural Volcani Center in Israel made the third request. He is a masters student investigating the degradation of pyrene in soil and needs the measurements for extraction experiments. The contacts we make via this route might turn into more involved collaborations or just end up being interesting people to interact with. Either way it benefits science.

4) Implicit requests by keyword searches: We know from our Sitemeter and Google Analytics the keywords that people use to find our site. Most of the time they find the solvent and solute combination that they are looking for but occasionally a combination arises that we have not yet tried. In fact we set up this page with all of the compounds and solvents in my lab to catch new combinations that we could run. This is how the request for phthalic acid in chloroform found its way on the dosol sheet. If someone or something searched for it there is a need for it to exist.

Ultimately, whether entered by human or machine agency, the requests on the dosol list are simply that - requests. It is still left to me to assign the priority in a logical way, allocating resources in proportion to their importance in the portfolio of projects that we are pursuing. One of those projects involves listening to what the global chemistry community is asking for and betting that it will be useful to answer.

Minggu, 14 Desember 2008

Crowds, Solubility and the Future of Organic Chemistry

This week I participated in a Social Media Day at NIST. During my talk I provided an overview of our current work in using Web2.0 tools for doing Open Notebook Science in fields related to chemical synthesis and drug discovery.

During my talks I generally try to place our work in context and give the audience a sense of where I see science evolving. I often start with the increasingly important role of openness and at some point follow up with this slide showing the shift of scientific communication from human-to-human to machine-to-machine. My position is that we are entering a middle phase of human-to-machine and machine-to-human communication. This is essentially what the semantic web (Web3.0) is all about and the social web (Web2.0) is the natural gateway.


Giving a talk at NIST was particularly meaningful for me as an organic chemist. This is an organization that has always been associated with authoritative and reliable measurements. In chemistry, many properties of compounds are deemed important enough to be measured and recorded in databases.

Given that the vast majority of organic chemistry reactions are carried out in non-aqueous solvents, isn't it surprising that the solubility of readily commercially available compounds in common solvents is not considered a basic property, like melting point or density? You won't routinely find these values in NIST databases or ChemSpider or even toll-access databases like Beilstein and SciFinder. Tim Bohinsky has reviewed the literature to provide an idea of what is available for a few classes of compounds.

I think that the reason for this lack of interest in solubility measurements relates to the way synthetic organic chemists have learned to think about their workflows. Generally, the researcher sets up an experiment with the intention of preparing and isolating a specific compound. The role of the solvent is usually just to solubilize the reactants - it is then commonly evaporated for chromatographic purification of the product. Even in combinatorial chemistry experiments where products are not purified for a rough screening, the expectation is that compounds of interest will be purified and characterized at some point.

The advantage of this approach is that it is relatively reliable. Column chromatography and HPLC may be time consuming and expensive but these purification techniques will work most of the time. However, they are difficult to scale up and are not environmentally friendly.

Sometimes, during the course of a synthesis, a compound crystallizes either from the reaction itself or by a recrystallization attempt. When this happens, it is a lucky day. The problem is that you can't routinely guarantee purification of compounds this way. In the academic labs where I worked, that was always the case, although there were rumors of gurus with magic hands that could get crystals more often than most. Before chromatography became widely adopted I am sure chemists were much more adept at recrystallization by necessity.


But now technology is allowing for different ways of thinking about organic chemistry. Instead of attempting to make a specific compound, why not think about making any compound that meets certain criteria? If the objective is to inhibit an enzyme, then docking or QSAR predictions would be the first criterion. But we can add to that the requirement for the compound to be made from cheap starting materials using convenient reaction conditions and that it be purifiable by crystallization.

This last requirement would be predictable if we had robust models for non-aqueous solubility. We can only do that if we gather enough solubility measurements - and that is the point of the Open Notebook Science solubility challenge. We want it to become as easy to look up or predict the solubility of any compound in any solvent at any temperature as it is to Google. For a taste of things to come, play with Rajarshi Guha's chemistry Google Spreadsheet that calls web services to calculate weights and volumes of compounds based only on the common name and number of desired millimoles.


In thinking about what the future will look like it is tempting to imagine complex extrapolations of current concepts caught in the first part of the hype cycle - nanotechnology and artificial intelligence are good recent examples.

But much of the real progress is reflected by a simplification that is almost invisible and underestimated in its power to change the way things are done. Blogs, wikis and RSS are a wonderful example of that. Technically, these are very simple software components - probably not something people in the 1980s would have predicted to be the "advanced technology" to explode in the 21rst century. But it is precisely that simplicity, coupled with reliable free hosted services, that accounts for the explosive adoption of these tools.

Similarly, as a graduate student in the late 1980's, I imagined complex chemical reactions in academia by now to be carried out routinely by robots and synthetic strategies to be designed by advanced AI. From a purely technical standpoint, academic research probably could have evolved in that way but it didn't. That vision simply was not a priority for funding agencies, researchers and companies.

But now the Open Science movement, fueled by near zero communication costs, can add diversity to the way research is done. I'm predicting it will favor simplification in organic chemistry. Here's why:

In a fully Open Crowdsourcing initiative, where all responses and requested tasks are made public in real time, the numbers will dictate what gets done. There will always be more people with access to minimal resources than people with access to the most well equipped labs. Thus contributions to the solution of a task will likely be dominated by clever use of simple technology. This is what we expect to see for the ONS solubility project. As long as competent judges are available to evaluate the contributions and strictly rely on proof, the quality of the generated dataset should remain high.

Note that this may not be the expected outcome for crowdsourcing projects where the responses are closed (e.g. Innocentive). Responses to RFP's from traditional funding agencies, where funds are allocated to specific groups before work is done, are also unlikely to yield simplicity. Even to be eligible to receive those funds generally requires being part of an institution with a sizeable infrastructure. I'm not saying that traditional funding will disappear - just that new mechanisms will sprout to fund Open Science. The sponsorship of our ONS challenge with prizes from Submeta and Nature is a good example.

For synthetic organic chemistry, it doesn't get much simpler than mix and filter. We've already shown that such simple workflows can be automated with relatively low-cost solutions, with the results posted in real time to the public. Add to this crowdsourcing of the modeling of these reactions and we start to approach the right side of the diagram at the top of this post. See Rajarshi's initial model of our solubility data to date to see how it adds one more piece to the puzzle.

Minggu, 28 September 2008

Open Notebook Science Challenge

The wiki for the Open Notebook Science Challenge that I proposed during my UK trip is now available. We are currently looking for sponsors and participants.

Open Notebook Science Challenge

What?

The first round of this challenge calls upon groups or individuals with access to materials and equipment to measure the solubility of compounds in organic solvents and report their findings using Open Notebook Science .

Why?

Understanding exactly how an experiment was performed is essential to the efficient progress of science. There are no absolute facts in the scientific literature; every measurement reported is only meaningful within the full context of how it was generated. The purpose of a laboratory notebook is to report as much of this context as is reasonable. But to find trends data must be abstracted to a level where they can be manipulated in tables and charts. This is not a problem as long as one can drill down from each data point in a chart to the full context found in the laboratory notebook.

For example, a Google search for "vanillin solubility in THF" pulls up a lab book page EXP207 where it is reported to be 3.89M. This number might be used in a table of someone trying to quantify trends or test a mathematical model, in which case reliability of the number is important. By reading the lab notebook page it becomes clear that 118.5 mg of solid was measured on a scale with 0.1 mg accuracy. However only one measurement was obtained. All kinds of other details which might be important are provided, for example how long the mixture was vortexed, at what temperature and the physical appearance after evaporation. If this number turns out to be an outlier, one can investigate if a calculation error was the cause by inspecting the linked spreadsheet.

However, if a researcher is simply looking for the feasibility of making up a 2M solution of vanillin in THF for a reaction the margin of acceptable error is so wide that the answer is almost certainly "yes".

The purpose of Open Notebook Science is to allow immediate communication of scientific results. The value of these results will depend upon the quality of the laboratory notebook and the linked raw data. Publication in peer-reviewed journals is still an extremely important part of this process but it is not an appropriate vehicle for the efficient communication of this type of information.

In fact, one of the motivations for participating in this project is that we will collect data from sufficiently well recorded experiments and publish them in a peer-reviewed journal with the participation of the researchers as co-authors. We aim to build a mathematical model to predict solubility using the results obtained from this project.

Who?

Organizers

Jean-Claude Bradley
Cameron Neylon
Rajarshi Guha (modeling)

How?

Simply request an account on this wiki and start recording experiments using a format similar to UC-EXP207 . The organizers will provide feedback in the form of comments in bold and italics directly on the wiki. Hitting the Recent Changes link on the left navigation bar is a good way to keep track of edits.

When?

Now!

Kamis, 13 Desember 2007

Chemistry Crowdsourcing with Open Notebook Science

I recently submitted a Letter of Intent for the NSF Cyber-Enabled Discovery and Innovation competition. Kevin Owens is a co-PI and will assist with the laboratory automation component. ChemSpider will contribute the database support. The pre-proposal is due in early January 2008 and we'll be writing it openly here. Comments are welcome.

We would ultimately like to enable the chemistry community to directly control the actions of a robot to help us understand some chemistry problems. As we make our way towards this goal, it would be very useful to start with suggestions for protocols to be executed by students we currently have in the group.

We already have a mechanism in UsefulChem to post experimental plans. In order to make the transition to full automation easier, it would be preferable if suggested protocols are even more specific than what we currently have listed. For example, instead of describing a general procedure like EXPLAN005, actually specify all of the compounds, amounts, mixing times, etc. This way the protocol can just be copied and pasted in the Procedure section of a new experiment, executed faithfully and reported in the main experiment list.

The main puzzle to solve is the prediction of which Ugi products will precipitate. A hypothesis might be that a precipitate will always occur from methanol at a certain minimal concentration of a certain reagent. Another approach might be based on the predicted molecular descriptors of the Ugi products. We might also start with as few assumptions as possible and use a genetic algorithm to evolve a solution. We'll be doing some of these but clearly there are more ways to solve this puzzle than we have resources or expertise.

So if anyone is interested in participating at this stage contact me to get access to the wiki and further discuss. Other examples of chemistry crowdsourcing : Chemmunity, The Synaptic Leap, OrgList, Chemists Without Borders and ChemUnPub.

Here is the LOI:

Chemistry Crowdsourcing using Open Notebook Science

The current system of dissemination of scientific data and knowledge is far less efficient than it needs to be to facilitate improved collaborative science, especially considering current publication vehicles and infrastructure. There is a growing movement promoting more Open Science, with the belief that a more transparent scientific process can perform far more effectively. The logical extension of this concept is full transparency - exposing a researcher's complete record of progress to the public in near real time. Not only will such a process enable ongoing data sharing it also provides an opportunity to develop collaborative communities of scientists and, at the conclusion of data acquisition, can enable communal extraction of conclusions when necessary. We have named this approach Open Notebook Science and have demonstrated its implementation and feasibility with the UsefulChem project, started in the summer of 2005, with the aim of synthesizing novel anti-malarial compounds. Our system currently uses free hosted services using general blog and wiki functions to facilitate replication across any scientific domains. These services are not chemically intelligent and are limited to text and graphic based data sharing only. For Open Notebook Chemistry the ability to intelligently manipulate, manage and search chemical structures and associated data is necessary and we have demonstrated proof of concept capabilities by integrating with the ChemSpider service, a free access online database managing chemical structures and focused on developing a structure centric community for chemists. This work will require the development of a chemically intelligent software platform to extend the capabilities of both the blog and the wiki environment for managing Open Notebook Science. The exposure of raw experimental procedures and data in a semantically rich format will enable the participation of both human and autonomous agents in the process of scientific discovery. This phenomenon of spontaneous group intelligence, referred to as "Crowdsourcing", has proven valuable in several contexts. Already, productive collaborations have been forged within the UsefulChem project with groups from Indiana University, Nanyang Technological University, the National Cancer Institute and UC San Francisco.