search

Tampilkan postingan dengan label open science. Tampilkan semua postingan
Tampilkan postingan dengan label open science. Tampilkan semua postingan

Kamis, 16 Juni 2011

Live Tweeting Haumea: the Open Science Ratchet at work?

Eugenie Samuel Reich just announced on the Nature NewsBlog that astronomer Mike Brown live-tweeted his observations of a transit of dwarf planet Haumea by its moon, Namaka.

About a year ago, I wrote about Mike Brown and the controversy about the discovery of Haumea stemming from a competitor's more aggressive data dissemination practice. In that post I speculated that we could expect accelerated data sharing over time due to the Open Science Ratchet, where the actions of scientists that are most open set the pace for everyone else working on that particular project, regardless of their views on how secretive science should be.

I don't know if Mike Brown has changed his views on data sharing - or if he has always felt this way but thought it was too risky until now. Either way, he certainly is taking the lead at this point to demonstrate how radical openness can be done in astronomy!


Kamis, 26 Agustus 2010

Open Notebook Science in Drug Discovery at Opal Event

I presented on "Open Notebook Science in Drug Discovery" on August 24, 2010 at a panel on Industry and Academia part of the Opal Event "Drug Discovery: Easing the Bottleneck".
I only had about 15 minutes to present so I could not go into much detail but I did want to highlight the most recent work Andrew Lang and I (also with Peter Li from ChemTaverna) carried out involving solubility prediction and web services. Most of the attendees were from industry and I appropriately used the recent GSK malaria data sharing to introduce the talk. It is clear that there is a role for Open Science in drug discovery and I think that industry involvement will continue to increase in this area.

My co-panelist Rathindra Bose from Ohio University presented on his group's development of a novel cancer treatment compound based on platinum. He made the point that academic research complements that from industry by being able to explore more speculative hypotheses. The dominant hypothesis for the mechanism of action of platinum based drugs is binding with DNA. By exploring alternative scenarios, his group found an active platinum drug that does not bind with DNA.

During the preceding session on the Emergence of Biologics in Drug Discovery, Albert Giovanella from the University of Pennsylvania School of Medicine gave a particularly enlightening talk about comparing biologics with small molecule drugs. Although biological drugs tend to have less toxicity, the overall cost to bring them to market is still quite high and their cost to the consumer may be so high as to limit their impact. It looks like it will not be generally easy to translate new biomedical knowledge to a widespread impact on human health.

Selasa, 01 Juni 2010

Use of ONS to protect Open Research: the case of the Ugi approach to Praziquantel

As we were collecting reactions from The Synaptic Leap for the Reaction Attempts project, Andrew Lang noticed that there might be a quick synthetic route to praziquantel via a Ugi reaction. I researched it further and found a paper (Kim et al 1998) where Ugi product 1 was indeed converted to racemic praziquantel via the Pictet-Spegler cyclization.


Using Beilstein Crossfire the only synthesis of 1 I found involves a multi-step amidation strategy. But this compound should be accessible in one step from commercially available starting materials via a Ugi reaction (shown above). Since all the starting materials are liquids we have some flexibility with solvent choice. Khalid first tried it in methanol EXP258 a few weeks ago but did not get a precipitate. He was going to monitor it by NMR next to see if the problem was high solubility of the Ugi product or with the reaction itself.

It was therefore with great interest that I read Mat Todd's report this morning on The Synaptic Leap that a German patent had been issued on this Ugi strategy to praziquantel. (TSL didn't provide a means of leaving a comment so I edited the page - which made me the author of that post but actually Mat wrote it)

I have often mentioned during my talks that Open Notebook Science could be used not only in a defensive manner to claim academic priority - but also as an offensive tactic to block patent applications. A company attempting to prevent the commercial exploitation of rival inventions has a few options. Where applicable, it can buy up an existing patent pool with the intention of sitting on it. For new inventions, it can do research and try to file patents before their competitors. But this is a costly process and it may make more sense to simply publish the inventions to create disclosed prior art, thereby blocking patent applications of their competitors.

But - as I and many others have discussed - the current publication system is not optimally suited for the purpose of simply disclosing and communicating science. Not only is it generally slow but the traditional article format requires a narrative of some sort - rarely can single experiments be published. This means that much (if not most) of research done by an individual or group will never be disclosed.

For these reasons I think that keeping an easily discoverable Open Notebook for projects designed to block patent submission by competitors makes a lot of sense - both economically and from a workflow perspective. Since researchers already have to keep a lab notebook, making it public doesn't impose the added time that writing an article or patent will require.

In this specific example of praziquantel we were too late. But if we had recorded this experiment a few years ago it might have worked to block Domling's patent. Now, it isn't clear to me that EXP258 would have been enough to do that. The strategy to make praziquantel via a Ugi reaction was clearly stated but the experiment was not conclusive. However, since Domling reported that methanol worked I am sure that we would have had the "reduced to practice" evidence in the notebook shortly.

Above I used a company as an example of a party motivated to disclose inventions to protect their interests. In our case it would not be a company but rather the entire Open Science community. It is in our best interest to keep our scientific territory as unencumbered by patents as possible. Keeping Open Notebooks might be one of the simplest means of ensuring that.

Consider a humanitarian organization that might want to manufacture praziquantel. I haven't researched it but presumably the Domling patent was filed in a number of countries beside Germany. In order to consider using the Ugi strategy, the organization would now have to deal with the patent holder. This might be the factor that makes this route untenable. Patents have proven to be problematic for humanitarian aid - even in the simple case of providing food.

But all is not lost. In addition to offering a simple 2-step synthesis of praziqantel, the Ugi route offers an easy way to make large libraries of analogs. Optimally we would like to work with someone who has experience with docking praziquantel. It might be interesting to screen not only the praziquantel analogs but also the uncyclized Ugi products themselves. When we did this for malarial enoyl reductase inhibitors (D-EXP005) we found that we did not need to cyclize to obtain compounds predicted to bind. This ultimately led to active compounds.

Sabtu, 27 Maret 2010

Education 2.0: Leveraging Collaborative Tools for Teaching

On March 25, 2010 I presented at the Drexel E-Learning 2.0 Conference on "Education 2.0: Leveraging Collaborative Tools for Teaching". It was an opportunity to update my slides with what I did and learned from the Chemical Information Retrieval course I taught over the Fall 2009 term.

I described using a wiki to organize course content and to allow students to contribute useful resources. Their assignments were also designed to be useful to other students in the class as well as to the general library and chemistry community.

I covered using wikis and other collaborative tools to mentor students doing laboratory research with Open Notebook Science. At the end I provided a quick overview of using games and Second Life for educational purposes.

Kamis, 04 Maret 2010

Nature Precedings as an Archiving Tool for ONS Solubility Book

The issue of archiving and citation is a topic that is usually raised whenever I give a talk about Open Notebook Science. We have recently tried to address this using several complementary strategies.

The publication of a book containing a snapshot of all the values obtained from the Open Notebook Science Solubility Challenge has turned out to be a convenient mechanism. By using LuLu, the book can be either downloaded for free as a PDF or ordered as a physical copy for just the printing and shipping charges.

However, Lulu does not have a convenient method of keeping track of different editions of the book and it is unclear how to best cite them.

Nature Precedings solves both of these problems quite nicely. I have uploaded the PDF of each book edition to NP and the versions are automatically linked to each other. In fact if you try to access an older edition, NP pops up a warning that a more recent version is available with the corresponding link (see image below).

Precedings also provides information about how to cite the document, including a DOI for each version. Unfortunately it appears that it can take some time for the DOIs to resolve. Links to different versions can also be formatted like this:
http://precedings.nature.com/documents/4243/version/1
http://precedings.nature.com/documents/4243/version/2
http://precedings.nature.com/documents/4243/version/3
Links to the Lulu version of each book are also provided, which is convenient for anyone who might want to order a physical copy.

At this time Precedings does not accept zip files containing the full archive of the source files for each book version - although a link to the archive is provided in the preface of the book. We have found that our library's DSpace repository is a convenient location for these.

Senin, 08 Februari 2010

Funding Agencies and Open Science

I've been invited to participate in a panel discussion on "New tools in research, teaching, and publishing" on May 24, 2010 at the annual PI meeting for the Integrative Graduate Education and Research Traineeship (IGERT) program at NSF. After speaking with program manager Vikram Jaswal, I feel encouraged that funding agencies are interested in exploring the emerging role of Open Science and related novel communication channels for facilitating scientific progress.


The role that funding agencies can play in Open Science has been the subject of some discussion in the blogosphere. One view is that they can require more openness as a condition of funding. The NIH's requirement to make papers resulting from funding Open Access after 12 months of publication is a step in that direction. There is a debate about whether this should be extended to Open Data - even to the point of Open Notebook Science, where even failed experiments would be shared for the scientific community to learn from.

I tend to prefer the carrot to the stick. I think that funding agencies could value plans for "sharing beyond the norms" in proposals without imposing strict requirements. In the long run OS will succeed because each stakeholder (researcher, funder, publisher, etc.) acts out of selfish motives. I believe that the most effective way to stimulate this selfishness is to show concrete examples of practice and benefits.

Funding agencies should see the benefits of OS as a higher ROI - in terms of knowledge gained and shared with the scientific community - as well as the wider population ultimately footing the bill. A perceived downside of higher transparency might be the greater difficulty in fueling hype cycles. Most things aren't as pretty up close and science is no exception. If you measure success as the absence of failure and ambiguity then increased transparency is going to be a problem. Most experiments are failures of some sort (as the saying goes - if you're not failing you're not trying hard enough). But failed or successful - both categories of results can be useful to others if they are made available in a way that they can be discovered easily. Funding agencies can help transparency by making it clear that the whole truth is more valuable than a subset of the truth presented in a way that might be conveniently misleading.

This doesn't mean that you can't put your best foot forward and give a slick PowerPoint presentation to guide your audience. It is ok to construct an easily digestible narrative of your research. It is ok to distill your work down to key conclusions. It isn't necessary to confuse your audience with every ambiguous result and unanswered question.

But - in addition to the streamlined version of your work - if you provide all the details of the failures and ambiguities for those who can benefit from further exploration of what you have done - there is a great potential for accelerating the scientific process. For a funding agency OS can mean a bigger bang for the buck.

Kamis, 15 Oktober 2009

NERM 09 session on Chemistry on the Web

Last week, on October 9, 2009 I presented at the ACS NERM conference. Martin Walker hosted a session on Publishing and Promoting Chemistry in the Internet Age. All of the talks were quite interesting and fit perfectly with the topic:
Martin Walker Chemistry on the Internet
Elizabeth Brown The Chemist's Toolkit for Publishing and Promoting Your Work On the Internet
Antony Williams Navigating the Complex Web of Chemistry Using ChemSpider
Jean-Claude Bradley Leveraging Transparency and Crowdsourcing in Chemistry Using Open Notebook Science
My talk consisted of an overview of Open Notebook Science with some new content on solubility prediction algorithms written by Andrew Lang and a few example of students taking a Chemical Information Retrieval class at Drexel University using research logs on a wiki to flesh out their projects.



Selasa, 18 Agustus 2009

Spectral Game talk at ACS Fall 09

Yesterday (August 17, 2009) I gave my talk on the Spectral Game at the Using Technology to Enhance Learning in Organic Chemistry symposium at the American Chemical Society meeting. I was not able to attend the entire symposium but luckily I did catch David Soulby's talk on using Google groups to distribute NMRs for labs that require many students to submit samples. I am a fan of using free and hosted services to simplify workflows of all types.

Also in attendance at the symposium were Liz Dorland and Bob Hanson. It was good to catch up with them. Bob shared a story of how he has been assigning his students tasks in his organic chemistry class which lead to updating Wikipedia. There is so much potential for using the educational infrastructure to create better scientific content for everyone.

My talk on the Spectral Game highlighted the role of openness in teaching and research to create new educational tools, especially for learning NMR. Tony Williams said a few words at the end about ChemSpider, RSC and some upcoming opportunities to publish synthesis articles on ChemSpider.

Kamis, 12 Februari 2009

Open Notebook Science, Reproducibility and Exclusion

There has been a fairly active conversation about Open Notebook Science over the past few days on FriendFeed. Some of the points I have wanted to make wont fit there so I'll post them here.

A short definition of ONS on Wikipedia: Open Notebook Science is the practice of making the entire primary record of a research project publicly available online as it is recorded.

I have repeatedly said that Open Notebook Science is probably not the best choice - at least initially - for most researchers interested in dipping their toes into Open Science. So why do I care so much about pursuing it?

It has to do with taking us past the tipping point to runaway real time open collaborative crowdsourcing in science.

There are two properties of ONS that at least set the stage for such a scenario.

1) Reproducibility: The primary purpose of a laboratory notebook is to make experiments reproducible for the researcher who recorded it. You can't improve on a process if don't keep a detailed record of what exactly happened in a given trial. A secondary purpose is to prove what a researcher knew and did at a specific point in time. This can be useful for patent enforcement if the notebook is kept private. There are other applications, but in general the sharing of experimental details is typically extremely limited in science.

I believe that this is an artificial barrier still standing to a large extent because of inertia. In the past, even if they wanted to, researchers could not realistically make this information public because of the absence of a convenient publication vehicle. But that is changing right now - technology is no longer the bottleneck. Very high quality free hosted services exist now that permit sharing with very little additional effort. All we have to do is record our laboratory notebook (which has to be kept anyway) on media that are easy to share and automatically indexed on major search engines.

By definition, the notebook should have all the information necessary to reproduce the results you obtained. If it is published in close to real time, someone who doesn't know you can read the details of an experiment you did today and contribute to the advancement of your project tomorrow. Or they may use your information for their own project. As long as they also maintain an Open Notebook knowledge can spread extremely rapidly. The efficiency of such a system between strangers is probably far greater than most scientific collaborations between researchers who already know each other.

2) Exclusion. If I come across an experiment you did yesterday and I have a desire to contribute meaningfully to your project, before executing the next experiment, I will want to look at all of the related experiments in your notebook. Of particular interest will be "failed experiments" to avoid repeating the same attempts. Or I may want to repeat one of your failed runs because I don't think that you properly controlled some parameter or made a mistake in the analysis.

The point of an Open Notebook is to get the truth - the whole truth about what you did and did not do. If I don't find what I am looking for in your notebook - and you have declared it to be an Open Notebook - then I can safely assume that you have not done it and I will feel confident to invest my resources to do that next experiment.

For an example of this system at work consider the Open Notebook Science solubility challenge. If I want to contribute to the project, the first thing to do is take a look at what has not been extensively measured yet. I can do this via the web query or directly on the experiment list page. By using the web query tool I can also look for contradictory results and try to resolve them. Or I may wish to include a control in my measurements and pick a compound that has been measured reproducibly using different techniques. I probably also want to look at the minute details of exactly how previous researchers applied their technique.

Now what would happen if we had adopted a Partial Open Notebook Science approach where we delayed recording lab notebook pages until a paper was published? Or what if we used the PONS variant of only recording experiments that "worked"?

Had we done anything short of a fully Open Notebook the project would have never gotten off the ground.

Minggu, 26 Oktober 2008

There are no facts: my position at NSF eChem workshop

I recently attended an NSF workshop on eChemistry: New Models for Scholarly Communication in Chemistry in Washington (Oct 23-24, 2008). The group consisted of about a dozen members, including publishers, social scientists, librarians and chemists. For background, this was the mandate:
Many scholarly communities have embraced new web-based models for disseminating the results of their research. These models include open access to formal publications and "gray literature", access to primary data and the tools to manipulate and visualize that data, interactive peer review, and integration with on-line discussion tools such as blogs and wikis. According to their advocates these new models make the scholarly process more transparent and substantially improve the opportunities for examination, re-use, and enhancement of new results.

This workshop will focus on Chemists who have generally been indifferent or resistant to these web-based models and to open access. By and large they continue to publish results in journals to which access is restricted to subscribers and reuse is limited by copyright. This lack of interest may have a number of origins including the different funding methods available to chemistry, the prevalence of industry participation and associated opportunities for profit from results, concerns about confidentiality and privacy, the possibility of longer term use of the data by their originators, or other aspects of the social and political organization of research in chemistry. The workshop will bring together experts from the chemistry, information science, open access, and science and technology studies communities to examine the multiple factors that influence adoption of new scholarly communication models.

The outcomes of the workshop will be reported in a white paper that will be made publicly available via this web site. The report will provide funding agencies, including the National Science Foundation and the JISC in the UK, with suggestions for targeted research programs that further examine the issues discussed at the workshop and that improve the communication and dissemination mechanisms that underlie chemistry scholarship (and internet-based scholarship in general).
Although the final report will be made publicly available in a few months, the presentation materials are not. After some discussion, I was permitted to liveblog the meeting under the Chatham house rule: Day 1, Day 2.

Of course individual participants may share their own presentations - here is mine. I can also share the scenario of the research process Jane Hunter typed up based on discussions from our sub-group between her, Jeremy Frey and myself.

My position statement and my main contribution to the workshop revolved around Open Notebook Science and its role in making the scientific process better through transparency. This is an extension of a statement I made a year ago on the importance of replacing trust with proof.

There are no facts in science - only measurement embedded within assumptions.

There are properties that have been determined so many times by different researchers and different techniques that we can treat a narrow range of values by consensus as if they were absolute facts. An example would be considering the boiling point of methanol at 1 atm to be 65C within one degree of accuracy. For most purposes that will suffice, as long as we understand the source of our confidence.

The problem arises when we treat rarely measured properties as facts simply because they are printed in peer-reviewed articles or tables in books. We teach our students not to trust numbers in Wikipedia but have no problem if they can cite a reference in a peer-reviewed journal, even without thoroughly analyzing the experimental sections.

We delude ourselves into thinking that we can appreciate our uncertainty of the value of a property simply by taking multiple measurements, taking an average and reporting standard deviation. That is actually a useful thing to do if we remember that we are measuring random errors and completely ignoring systematic errors, which are possibly very common in infrequently measured properties.

What is the solubility of 4-chlorobenzaldehyde in chloroform? UsefulChem experiment EXP208 reports it to be 0.07 molar. It was measured only once but I think duplicate runs would have come out pretty close to that. It might have slipped under the radar if it had not been measured in parallel with other chemically similar aromatic aldehydes with values all much greater than 1 molar. It just didn't make sense so we looked at the conditions reported in the experiment and the boiling points of all the compounds - this one had the lowest value (214 C at 1 atm). The pressure had not been recorded during the course of the experiment but when empty the Speed-Vac could go as low as 0.1 Torr, which would reduce the boiling point close to room temperature.

The next most volatile compound in this group was 2,6-dichlorobenzaldehyde. It was calculated by ChemSpider to be 239C at 1 atm, which is reasonable based on the 4-chloro analog. But here's an interesting twist - the reported boiling point is 165C on this MSDS sheet. It should be simple enough to see if that is an error by clicking through to the lab notebook page that generated that MSDS sheet... oh wait... MSDS sheets don't require proof, just this handy disclaimer: "We have not verified this information, and cannot guarantee that it is up-to-date." It also looks mighty trustworthy: "the page is maintained by the Safety Officer in Physical Chemistry at Oxford University". I'm not knocking Oxford - this is standard practice for the flow of chemical information in the current culture.

The bottom line is that 2,6-dichlorobenzaldehyde didn't evaporate off - we get a value of 3.4 M in chloroform. Now is it possible that some of it evaporated under the conditions of that experiment? Maybe but it my call that we're going to use that number for now as a good enough approximation for our model. It is possible that your application might have a different requirement. At least you have the information available in the Open Lab Notebook to make the call.

The solubility of 4-chlorobenzaldehyde in chloroform was measured again, this time monitoring the pressure and minimizing time on the Speed-Vac. The pressure varied over the course of the evaporation, making it impossible to neatly summarize in the experimental section of a paper. The measurement was done in duplicate in EXP209 and comes out at 3.61 molar with a standard deviation of 0.02. That isn't a fact but a good enough number under these circumstances to pretend it is and use it for our model. We'll see how it plays out when we have different researchers and use different techniques.


Jumat, 05 September 2008

Mid UK Open Science Trip Report

I’m about half way through the UK trip that Cameron Neylon organized for me. So far it has proven to be a productive though exhausting Open Science fest. First I stopped by UKOLN at the University of Bath on August 29. Then the Nature Science Blogging conference in London followed by an Open Science Workshop at Southampton University then back to London for a talk at the Nature offices.

Cameron was with me at all these events and we even co-presented at Nature. It was a special pleasure to be able to report on the results we obtained in his lab just the previous day. Now a Google search for “boc-glycine solubility THF” pulls up the UsefulChem lab notebook page EXP207 as the first hit. What better example of Open Notebook Science in action?

The FriendFeed effect was in full force at a few of these events. People are finding it natural to microblog sessions by posting comments. Richard Grant, Cameron and I made good use of it at the Southampton conference. When a speaker started to present someone would start a thread and others would comment on that thread. It is interesting to see how FriendFeed is evolving within scientific communities. Some perceive that it is mainly to be used for ephemeral conversations. But why not bookmark conversations to document conferences for the longer term? A big problem is FF doesn’t generally make its way into the Google index but bookmarking should force that to happen for selected content.

The Southampton conference was also a great opportunity to discuss in detail the challenges and opportunities facing Open Science. If the ideas discussed there get fleshed out in document we may end up referring to a “Southampton Resolution on Open Science”. By the end I think we came to an agreement that a viable compromise might consist of sharing all raw data files associated with a paper after publication. This is a far cry from the ideal of Open Notebook Science, aiming for “no insider information”, that Cameron and I advocate. But it is a step in the right direction and it is not unreasonable to try to get a number of scientists on board.

I really enjoyed my conversation with David De Roure from MyExperiment at Southampton U. There is great potential for bringing cheminformatics workflows up to speed with bioinformatics. I agree with David that using ChemSpider web services on Taverna might bring us a long way in that direction. In organic chemistry automatically calculating masses and volumes of reagents based on moles and chemical identifiers like SMILES or InChIKeys certainly would come in handy for minimizing errors and planning time.

After the London Science Blogging conference, Cameron, Egon and I headed out to Peter Murray-Rust’s house for lunch and coding. We created a CMLreact file for one of the Ugi experiments that was part of our paper about to appear in JoVE (see here for Precedings version). This can be approached from many different levels of abstraction. We chose to focus on the equimolar 0.4M experiment run in triplicate and ignore the specific sequence or manner of reagent addition. As long as the abstractions point back to the original lab notebook pages with full details, simplifying the representations of reactions is not a problem. But there is certainly not a standard way of this as of yet using CML.

I'll put more up shortly, including my presentation in Manchester. Right now Duncan Hull is waiting for me to join him at the pub.....

Selasa, 26 Agustus 2008

Happy Accidents: A Must-Read for Open Scientists

I usually limit my book reviews to Goodreads or Shelfari but this one deserves much more attention.

In Happy Accidents: Serendipity in Modern Medical Breakthroughs; When Scientists Find What They're NOT Looking for, Morton Meyers reviews examples of the unpredictability of scientific progress.

This could just be a collection of interesting anecdotes - and some of the stories are truly fascinating. My favorite is probably the discovery of platinum compounds for the treatment of cancer. It came about from the accidental electro-dissolution of a platinum electrode during an experiment studying the effect of electricity on cell cultures!

But Meyers goes further and uses these examples to make larger observations about the way science operates today in both academia and industry. A quote from the preface foreshadows the tone of the book:
The dominant convention of all scientific writing is to present discoveries as rationally driven and to let the facts speak for themselves. This humble ideal has succeeded in making scientists look as if they never make errors, that they straightforwardly answer every question they investigate. It banishes any hint of blunders and surprises along the way. Consequently, not only the general public but the scientific community itself is unaware of the vast role of serendipity in medical research. Typically, a discoverer may finally admit this only towards the end of his or her career, after the awards have been received.
And starting on page 304:
An applicant for a research grant is expected to have a clearly defined program for a period of three to five years. Implicit is the assumption that nothing unforeseen will be discovered during that time and, even if something were, it would not cause distraction from the approved line of research. Yet the reality is that many medical discoveries were made by researchers working on the basis of a fallacious hypothesis that led them down an unexpected fortuitous path.
....
The peer review system forces investigators to work on problems others think are important and to describe the work in a way that convinces the reviewers that results will be obtained. This is precisely what prevents funded work from being highly preliminary, speculative or radical. How can a venture into the unknown offer predictability of results?(my emphasis)
....
Indeed the basic process of peer review demands conformity of thinking and disdains a maverick's approach.
....
What it comes down to is this: Who on a review committee is the peer of a maverick? (my emphasis)
The fact that some of us in the Open Science community are discussing this does not mean that we are advocating for the abolition of peer review or the NIH. We are not that naive. We still submit proposals and manuscripts for publication in peer-reviewed journals (although given a choice we probably would pick an Open Access journal over one running on a paid subscription model).

The point is what we do in addition to all those traditional processes.

We can share our failed experiments. We can share our research plans. We can discuss science freely admitting what we don't know. We can record our talks at closed meetings and make them public. We can initiate and participate in serious scientific conversations going on in the blogosphere without worrying about everyone's title and rank.

Basically, we can collaborate in ways that are most conducive to serendipitous discoveries. The free social software, databases and other infrastructure now available make this information exchange easier than ever.

The key question for a researcher today: to hoard or not to hoard?

To me, it seems likely that data hoarders will find it more and more difficult to claim priority for a contribution when competing against loose associations of open collaborators motivated by insatiable curiosity.

Some of the folks from the funding side are getting it. Take a look at SubMeta.

Rabu, 06 Agustus 2008

Scribd as a Repository for Proposals and Science Docs

Since Nature Precedings has tightened its policies and no longer accepts proposals, I have been looking for some alternative PDF repositories.

I think it is very important to have a convenient way to cite these documents on platforms with third-party timestamps and all the bells and whistles of web2.0 - ratings, comments, easy sharing tools, etc. Open Science is not just about what we are doing but also where we're headed.

Scribd seems to be a good solution. I've posted my last proposal to the Gates foundation there, in addition to the SCIEnCE site.

Jumat, 01 Agustus 2008

The BCCE, research discussions and good friends

I've just returned from the Biennial Conference on Chemical Education 2008.

I wasn't able to record my first talk on Communicating Results from Undergraduate Research because my computer crashed. However, I repeated many of the same concepts in my second talk on Open Notebook Science and Cheminformatics.

It was very fortunate that the BCCE was held in Bloomington at Indiana University this year because it was a great opportunity for me to meet up with Rajarshi Guha, David Wild and Amar Flood.

Amar and I discussed Open Notebook Science with his lab people and it may make sense to do this for some of their projects. We set up a wiki to explore that possibility.

Rajarshi and I discussed at length our collaboration on the prediction of Ugi precipitates and docking against falcipain-2. It is certainly easier to pour over the relevant papers and online documents when face to face.

These are the outcomes:

1) We're going to separate the problem of predicting the solubility of Ugi products from the problem of generating the best enzyme inhibitors. Our initial plan was to try to make the top ranked Ugi products for a given enzyme and hope that we generate enough precipitates from those results for Rajarshi to model. We're just not getting enough positive results that way to generate a reliable model in a reasonable amount of time. To stack the deck in our favor we're going to start with a well behaved Ugi reaction (EXP099) and modify the reagents one at a time.

2) We're going to separate the performance of the Ugi reaction from the solubility of Ugi products in various solvents. Once we have Ugi products in hand in pure form from a reaction in methanol we will simply measure their solubility in other solvents, starting with ethanol, acetonitrile, THF and toluene. Low solubility will not guarantee that the Ugi reaction will proceed smoothly to produce a precipitate when carried out in a given solvent but it will certainly be a great starting point.

3) We're going to actually measure the solubility instead of just noting soluble or insoluble, as we have been doing in our Ugi master table. This will make it much easier for Rajarshi to come up with a robust model. Kevin Owens is already set up with a SpeedVac in his lab and that will help immensely. Rajarshi made the point that models predicting solubility in non-aqueous systems are needed and could be quite helpful to the chemistry community. He will be using the crystallographic data of our precipitates in his calculations. More on this later...

4) We're going to require the reactions to be easily amenable to automation, even if we can carry them out manually sometimes. For example, some of the reactions had starting materials that were not very soluble in methanol by themselves, even though they went into solution when combined with the other starting materials and then generated a product. This is interesting behavior but extremely inconvenient for automation because we can't make up stock solutions of reagents at 2M concentration in methanol.

5) We're going to require the reactions to be fast. Some reactions required several days to complete. These will now be considered to be negative for precipitation at the 16 hour mark.

Kamis, 10 Juli 2008

How should Open Notebook Science be used?

Maxine Clarke highlights a bit of recent controversy regarding Open Notebook Science that has been bouncing around the blogosphere and FriendFeedosphere.

There are some who interpret the ongoing publication of our laboratory notebook as an expectation for the world to read it like a magazine. For someone who is not a collaborator or working in a related area that would make about as much sense as reading the phone book.

Here is an example of how an Open Notebook should be used:

Based on our Sitemeter records, this morning someone from the UK seached for "purification of phenylacetaldehyde" on Google. The first hit was the UsefulChem experiment EXP037 performed by undergrad student James Giammarco on October 6, 2006.

This happens to be a "failed experiment". But here is the information that can be obtained from the wiki:

1) The picture of the experimental set up:



2) The results provide an NMR of the first distillate, showing exactly the nature and amount of impurities. This is before we started using JSpecView systematically so this spectrum is not interactive, it is just an image.



3) The discussion addresses a major problem with the first distillate and a suggestion is made about what might be the problem:
Although the HMR of 37-F1 shows reasonable purity for use as a reagent, it appears contaminated with droplets floating and sticking to the sides of the flask. This may be water condensed by the distillation in open air. Future distillation should be done under nitrogen.
Phenylacetaldehyde is reported to boil at 195C. Since 37-F1 was collected at 162C (approx), perhaps this corresponds to an azeotrope with water and would explain the contamination in the first fraction. In the future more careful measurement of the Ubergang temperature over the course of the distillation would help to clarify this issue.
4) The history of this wiki page reveals who made each contribution. A major advantage of using a hosted wiki as a lab notebook is that a reference to every version of the notebook can be made, marked with a third party time stamp:

On October 27, 2006 James recorded his log, the NMR results and a conclusion.
On October 29, 2006 I added some questions in bold, re-interpreted the results, added a detailed discussion and changed the conclusion.
On January 3, 2007 James corrected a spelling mistake I made.
On January 9, 2007 my graduate student Khalid added a link for phenylacetaldehyde.
On January 23, 2007 Khalid added InChI tags.
On November 22, 2007 James answered my questions in the log.

5) This is the kind of information that someone experiencing problems with purifying phenylacetaldehyde could use. In the experimental section of a chemistry article the most you are likely to get is "phenylacetaldehyde was purified by distillation", with no proof as to the actual purity.

Notice that the experiment is never really "completed". There is always more that someone could add in terms of the discussion or add a question anywhere in the document. That is how science actually happens for the most part and we can forget that when listening to a slick presentation at a conference or read a well-written article making it look like the project unfolded effortlessly from a perfect plan.

Coming back to the person who searched for and found this page. This lab notebook page may or may not help solve their problem directly. The most valuable information is probably the contact information for the people who contributed to the execution and analysis of the experiment.

We aren't hiding behind pseudonyms. We use our real names because we want to fully participate with the research community and we stand behind what we report. Shoot us an email for help if the notebook isn't enough.
Zemanta Pixie

Rabu, 09 Juli 2008

InkSpot and Open Drug Discovery

I had a nice long chat with David Leahy from InkSpot yesterday. There is actually a good discussion in the comments of David's most recent post.

What I gathered from our conversation is that InkSpot will be providing transparency on the computational side of the drug discovery process. This is something that could be very valuable to the Open Science community in general and to our group in particular as we look for new anti-malarial agents.

We've tried to record the docking results we've obtained from Rajarshi Guha using a wet-lab style report (for example D-EXP014). Hopefully there is enough information there for the docking experiment to be replicated but clearly it would be better if a workflow system were in place to automatically record the steps and make those records public and easily searchable.

David said he would start with a QSAR-style exploration of our Ugi product precipitate prediction problem. It will be interesting to see how this plays out.



Zemanta Pixie

Senin, 30 Juni 2008

Back from San Diego

Update: the presentation videos are now available here.

I'm back from the UCSD workshop on New Communication Channels for Biology held June 27-28, 2008.

The first day mainly consisted of talks while the second was very heavy on breakout sessions (see agenda). Probably the most beneficial part for me was seeing old friends or meeting in person several people I had only interacted with online previously. For example was nice to finally meet Dan Gezelter from the OpenScience project and Hilary Spencer from Nature Precedings. I also had a blast over enchiladas on Friday night with Mike Nieslen and Jen Dodd. Mike gave a very engaging talk on the Future of Science on Thursday night.

The presentations were recorded and should be available within a week on the wiki link above - I'll post an update here when I am made aware of it.

The take-home message for me resonated around the idea that there exists a group of really strong advocates for the Open Science movement. The actual tools, whether they be wikis, blogs, custom databases or open-source software, were secondary to the fundamental philosophy of openness.

The SciBarCamp breakout sessions were certainly very lively. Since we were mainly preaching to the choir, controversy was at a minimum. However we did acknowledge the difficulty in pursuing an Open Science agenda within the current establishment.

It was hard to resist pointing out the remarkable recent success of FriendFeed in facilitating communication between life scientists. Of course Deepak brought that up during his session.

The fact that this conference was organized and funded on such short notice is a testament to the commitment of the core Open Science group - thanks to John Wooley and
Srikrishna Subramanian for making it happen!


Kamis, 17 April 2008

Cell Article on Open Drug Discovery

Seema Singh wrote a review "India Takes an Open Source Approach to Drug Discovery" which just appeared in Cell: Volume 133, Issue 2, 18 April 2008, Pages 201-203. (The doi doesn't work yet but try this link in the meantime). You'll need a subscription to view it, an increasingly familiar irony of much of the Open Science discussion these days.

UsefulChem and our collaborators got a nice mention:
A related initiative is UsefulChem (http://usefulchem.wikispaces.com/), set up by Drexel University chemist Jean-Claude Bradley. Bradley has pioneered Open Notebook Science in which lab notebooks and raw research data are posted on the web for anyone to see and respond to (http://usefulchem.wikispaces.com/All+Reactions). As for success, Bradley says, “Probably the best example of a positive outcome from UsefulChem is finding two compounds that are somewhat active against malaria [in vitro],” blocking the activity of falcipain-2, a Plasmodium falciparum cysteine protease. “This demonstrates that a team of researchers can work together in the open—Rajarshi Guha from Indiana University did the docking calculations, my group at Drexel did the syntheses and Phil Rosenthal's group at UCSF did the testing.”

Minggu, 06 April 2008

Attila Csordas writing his thesis on a blog

Attila is writing his thesis openly and is welcoming comments:

From now on I start every “thesis live” post with the standard introduction: In the live thesis building blogxperiment I edit (digest, compile, write, rewrite, delete) my ongoing doctoral thesis in blog posts and put the parts together on thesis live. The title: The physiologic role of stem cells in tissues with different regenerative potential

I am not aiming any perfection, my focus is clearly on getting things (the PhD) done here. Anyway, I found the idea of “writing” a complete, lengthy and formal thesis outdated and inefficient (after all, scientists should conduct nice experiments and publish their results in short, inforich and accessible research papers in order to share it ASAP with the research community, not in book-length, otherwise unaccessible PDFs) and so I try to keep myself motivated by

- doing this “thesis live” series as an open science experiment and getting useful feedback from my fellow scientists and readers

- trying to include as many systemic, whole body level material into it that could be relevant for systemic regmed approaches

- reminding myself every day that without a PhD it is hard to move further in science officially (that’s the least motivating factor though as it is official)

Minggu, 23 Maret 2008

Expanding the UsefulChem Collaboration to Teaching Labs

A few weeks ago I received a very interesting email from Brent Friesen at Dominican University. He mused:
I am trying to put together a bridge between the type of opensource research you are instigating and the traditional Sophomore Organic Chemistry lab. There are over 4,000 college and universities in the United States - all of them teach Sophomore Organic Chemistry lab. How can we harness this resource? .....

SOC labs must fulfill 4 criteria:
1) inexpensive reagents and equipment
2) fit into the time constraints of 1 3-hour period per week.
3) Must be a robust reaction with fairly stable products. It doesn’t have to be “foolproof” but that helps.
4) Compatible with equipment, glassware, procedures that student know how to use and do.

Bottom Line: I am definitely interested in developing collaborative projects, especially if they can be performed as part of a Sophomore Organic Chemistry laboratory curriculum.
This is extremely encouraging news for open scientific collaborations and I am very impressed with Brent's initiative! It certainly is more work for him compared to maintaining the status quo.

Kevin Owens and I have discussed this possibility for some time now and he is willing to contribute by carrying out mass spectrometry if required.

Brent and I further discussed the applicability of the Ugi reaction that we perform in my lab because of its simplicity - 4 components are mixed together in methanol at room temperature and a Ugi product often precipitates within days, requiring only filtering to isolate. (see EXP150 for a good example)

However, one major limitation of the Ugi reaction for a teaching lab is the terrible stench of most isonitriles, one of the four key components. One way around this is to use isonitriles that don't stink, such as TOSMIC.

So as a starting point, we would like to start with a Ugi reaction which involves TOSMIC and has lead to a precipitate in our lab. There is only one example so far: 171H.



According to the lab notebook, all starting materials dissolved easily and a precipitate appeared after 2 days. The precipitate has not yet been isolated and characterized. Hopefully it will prove to be pure Ugi product, as other similar Ugi reactions have done.

It turns out that many of the top ranking compounds from Rajarshi's falcipain-2 docking run (V2) contain TOSMIC as the isonitrile. Phil Rosenthal at UCSF is still up for testing compounds for anti-malarial activity. Wouldn't it be a testament to the power of open collaborative science if a decent anti-malarial lead was uncovered through the routine teaching of undergraduate organic chemistry labs?

At the very least I'll bet it would be rewarding for the students involved.

Brent has placed the orders from Sigma-Aldrich and is moving full steam ahead. He recently wrote to me:
You know, I'm ready to dance and you are the only dance partner who seems to be ready and willing. Let's figure out a way to adapt the Ugi reaction to Sophomore Organic Chemistry laboratory and give it a try!

I would like to plan it for the week of April 7 and the following week...
Does not give us much time, but it can be done.
Keep track of Brent's activities on his blog.

Tags
171H InChIKey XRHVBZVCUGMHPN-UHFFFAOYAR
TOSMIC InChIKey CFOAUYCPAUGDFF-UHFFFAOYAC InChI=1/C9H9NO2S/c1-8-3-5-9(6-4-8)13(11,12)7-10-2/h3-6H,7H2,1H3