Friday, December 2, 2011

Quick summary ARL / DLF E-Science Institute Capstone -- Atlanta

Waiting for flight home from Atlanta from the ARL / DLF E-Science Institute Capstone event.  Overall, it was a very productive event, especially for discussions with Rob Olendorf (my collaborator and a data management librarian at UNM) and Dale Hendrickson (head of Library IT at UNM).  Almost all of the attendees were library personnel, and I learned a lot from my interactions and the ideas presented.  I thought I would jot down some ideas and action items.

First, action items.  We were encouraged to develop "next steps" for when we return to our institutions.  Here are some of ours:

1.  Incorporate Library interactions with the undergraduate physics course (PHYC 308L, electronics lab) I am teaching next semester.  This is a new course for me, and I won't have time or familiarity to diverge much from the very good plan that prior instructors have developed.  But I know enough that Rob, Dale, and I came up with some concrete ideas that will be great for spurring data management at UNM and with these budding scientists:

  • Guest lecture by Rob to describe data management and related library services.  I think this would be best for the second lecture period in the course.  Rob will describe issues of data management and we will announce our intention to integrate library data management into the course (below).  Rob will also give an overview of github and a quick "how to."
  • A substantial part of the course (as I understand from talking to prior instructors and students) involves developing LabVIEW code for circuit design and simulation.  I'm guessing (pretty sure) that no source code control or versioning is used.  I think this presents a good (not perfect) opportunity to teach the students how to use github for versioning and source code sharing.  I'm thinking it will be an integrated requirement for all of the coding during the semester.  The reason it's not perfect is because LabVIEW uses binary files, so some of the forking and merging functionality will not be appreciated.  Many of the students are experienced in Matlab, though, and where possible I will encourage moving to that platform.  Regardless of how this plays out, I think for sure the students will come away from the course with a fundamental knowledge of github and how wonderful it is for protecting and sharing code.  I think I will also require LaTeX for their final reports, which will work well with github.
  • Incorporate data management, using the Library Institutional Repository.  Some infrastructure and coordination with the library will be necessary here, because I don't think we've done it before at UNM.  Dale's idea is to create a "community" in the d-space IR for our course, e.g. "Junior Lab 308L."  The students will be in charge up uploading their final data sets (testing their circuits) into permanent, curated objects in the IR.  There may be difficulties with this, but I am confident that the students will come away with a good appreciation of the power of good data management, and, hopefully a real, curated data set as part of their career portfolio.
2. Participation is data management "group meeting."  The library currently has some kind of regular meeting like this, and I will visit one of their upcoming meetings.

3. (Mostly for Rob)--"finish" our pilot data management project.  Rob has been working on this for a long time and it hasn't been easy.  He is working on curating and archiving one of Andy Maloney's complete kinesin gliding assay data sets.  The uncurated data can be seen on our server.  I don't really understand how Rob is doing this, but he's done a lot of coding and is close to putting a curated version of that data set into our institutional repository.  There are 500,000 images in the set, and I think Rob said that involves more than 50 million lines of (XML?) code to describe it.  I may be getting terminology and numbers wrong ("schema," etc.) but the point is Rob is writing a lot of code to do it "right."  A finished product will serve as a great example to everyone on campus (and even broader), especially researchers as to what the library can provide for data management.  I think this will be a huge step for us at UNM and in convincing more researchers to collaborate with the Library for research data management.

There were many more "next steps," but they aren't coming to mind now.  More than just next steps, there were a lot of visionary ideas presented by groups at the capstone event.  Here are some that stuck in my mind:

1.  Graduate students are key to connecting data management librarians with research groups.  What seemed the best idea to emerge was that an existing pipeline to graduate students is the general requirement for "ethics / responsible research conduct" courses as part of NIH/NSF training grants.  Good data management is often part of these courses, and in my mind is essential for responsible / ethical research.  Given how these courses are usually implemented, I think it would be fairly easy for data management librarians to obtain one or more time slots to discuss data management with the graduate students.  Best would be "hands-on" coursework, where the students are asked to bring data to the course.  This was discussed a bit on a friendfeed thread.

2.  Our institutional group and at least one other (can't remember the institution) more than once mentioned a vision for the library providing more than just data curation / preservation / storage.  I don't have a good term to capture this area, but it involves capturing / helping with workflow (especially custom software used in labs for data management / processing) and data visualization.  In my mind, a ripe area for connecting with researchers is to work backwards from the traditional publication.  Currently, many libraries have an institutional repository that allows researchers to post PDFs of research papers.  And usually that's about it (from what I can see).  Working back upstream, what I think would be very useful is to provide a computational workspace (through the libary) where researchers can process and produce the figures in those papers.  As an example, my graduate student logs into the library workspace, and uploads the data needed to produce the final figures. The graduate student and me then use software on that workspace (maybe R, Matlab, Excel) to create the figures for the paper.  There is a versioning system to keep track of the code used to process the figures and the many versions created.  When the paper is submitted for peer review (the current standard), it is seamless to link each figure to the data sets and the code used to generate those figures, using either permanent URLs or DOIs.  For me as a researcher, I would LOVE such a system.  And talking with Dale and Rob, it doesn't seem too much of a pipe dream.  It's a lot of work, but I think it would be a huge step and improvement in data management and data sharing in research.  Successful implementation would also be a really great way to recruit more researchers into data management partnerships with the library.  An important component of this I forgot to describe above is that there will be experts in the Library (such as Rob) who can work side-by-side (virtually) with us to develop the data visualization code and figures.

3.  Related to item 1 above, I think connections with graduate students could be greatly accelerated by a grants / data management competition.  A $1000 dollar research grant prize, directly to graduate students for "the best data management," would I think be very effective.  Compared to what we need to accomplish to transform research and the library's involvement, $1000 every so often is not a lot.  But it would mean a lot to the graduate students in the competition.

4. The NSF Data Management Plan (DMP) requirement has already done a lot to connect researchers with data management librarians.  Rob estimates more than 30 faculty connections have been made for him at UNM because of DMPs.  I think this is just one great outcome of the DMP requirement.  And it illuminates a huge opportunity that I see for researchers and libraries.  In my specific case, if I get tenure at UNM, I want to pursue a couple training grants.  One specifically I would like to try for is an "open science" NSF REU program.  REU is "research experience for undergraduates," usually involving summer research internships for undergraduates from other institutions around the country.  I think an REU proposal with a heavy focus on "open science" and advanced data management would look very appealing to the NSF.  Of course I also think it would be very effective in training the next generation of researchers.  Importantly, though, I would need a lot of help to write this grant.  The Library's experience with DMP's can be extended to this effort and people like Rob and others will be essential in planning, writing, and executing the grant.  Moreover, I think other people on campus who are planning other training grants would get a big "broader impacts" boost from this kind of data management or "open science" collaboration with the library.  So, hopefully, our Research office can help coordinate these endeavors.

Many, many more ideas but I think I'm out of steam for now.  Overall, a great conference and I'm excited for pursuing these ideas!

Tuesday, November 22, 2011

An idea for wealthy donors: alternative to direct research funding: fund libraries to help with e-research

Next week, I am attending the E-Science Institute Capstone event, along with Rob Olendorf and Dale Hendrickson from U. New Mexico Libraries.  As part of our preparation for this event, we are interviewing several people around the university to capture their views on e-research.  Today, Rob and I interviewed Martha Bedard, Dean of the UNM Libraries.  Rob and Dale figured it would be good to have me lead the interview, since I'm coming from outside of the library and thus would ask different questions.  At least from my perspective, this was a success and I learned a lot in the generous one hour of time that Martha gave us.

At this point, I can't share the interview notes publicly, but I did want to share one idea that emerged during our discussion (and there were several good ideas!).  I'm having trouble getting the idea in writing so maybe by poorly blogging it, someone else can turn it into a good idea, if it's sensible at all.  Here's what I'm thinking: wealthy donors, or a group of donors that want to make a big impact on research at their university have at least the following two choices:

1.  Provide substantial money to fund research in a specific field, for example by providing 10's million dollars to fund a nanomedicine research center.  Or to build a new biomedical engineering building.  Etc.

2.  Provide substantial money (say $10 million) to the university library in order to vastly improve the ability of ALL researchers at the university to conduct e-research.  The money would go towards hiring many new library faculty and staff members and procuring and implementing storage and networking infrastructure.  The goal would be a completely transformed library that would make it easy and almost automatic for all university researchers to conduct connected, networked, open, archived, discoverable, etc. research.

Option 1 is common and makes a big impact on specific research fields.  Performing research in excellent facilities, with dependable funding is a great thing for researchers.  As far as I know, option 2 is less common, and I'm not aware of a good example.  But I think there'd be tremendous leverage compared to option 1.  The reason there is so much leverage is because currently the huge potential of "e-research" remains almost untapped.  There are shining examples of successes.  (For an excellent overview of the successes and the vast, untapped potential, read Michael Nielsen's excellent book.)  But in reality, for most researchers it's really difficult to manage data, share data, provide open access publications, etc.  And this is true even for researchers like me, who've decided to be as open as possible yet are finding it difficult to do so effectively!  So, it's basically true that there are huge technical barriers for most of the researchers to maximize the impact of their research by sharing.  Because we're so bad at it and because it's so difficult, I think there's a ton of room to make a huge impact at a university with a medium-sized grant.  I think the uinversity library is the natural and only choice to lead the effort.  And by doing so, it would impact all of the researchers across all of the disciplines (humanities, science, medicine, etc.).  How would they implement option #2?  I don't actually know, and that's a big reason why I want the library to do it!  Rob Olendorf, my collaborator at UNM on open data projects has a vision for how to make it seamless and almost automatic for researchers like me to connect, archive, and share our research and data.  I don't understand how that can work, and I don't have time to understand.  But I would LOVE to participate in that system.

That's the final key to the idea.  I think a university would gain a huge competitive advantage by becoming the "e-research leader."  There is a perception that most researchers are content with limited sharing and the status quo.  This may or may not be true.  But regardless, it looks like there is a lot of momentum, driven by the public interest, for funding agencies to go much further with data sharing, data management, open data mandates.  These mandates are scary to many researchers.  Even if researchers want to have excellent data management and share their data, it's almost impossible to do so now.  So, compliance will be a huge and new headache for researchers.  If a university could boast that compliance is "seamless and easy" it would be a real and strong recruitment incentive.  This probably sounds questionable to some, but I really see it as a huge incentive.  It would be just as appealing as the opportunity to work in a fancy new research facility.

Thursday, November 3, 2011

The inevitable spread of open science

Two things have happened this week that make me really happy about the research in our lab and the spread of open science.  First, we have a new undergraduate REU student, Alex Haddad, who has started her own open notebook science under the mentorship of Anthony Salvagno.  Her notebook is on wordpress.com and can be found here.  This is Alex's first experience in a research lab and she has immediately embraced open notebook science and she is excited about it.  One cool thing that I've noticed already is that her notebook entries are automatically linked in Anthony's notebook when she links to them.  Some kind of trackback thingy that I don't understand, but is great as far as good notebooks go.  An example can be found in Anthony's notebook entry, which automatically links to Alex's entry providing more information (see the trackback at the bottom of the page).  Welcome, Alex, to open notebook science!

The second thing that happened is that our former PhD student, Andy Maloney, just started a new postdoc at UT-Austin with Hugh Smyth.  This is going to be a very productive experience for both Andy and Hugh's lab, I am confident.  Most excitingly, though, is that Andy and Hugh have decided to incorporate open science into their projects!  I think this is very big news and a success for the spread of open science.  Major props to both Andy and Hugh for their willingness to carry out major parts of their research using open science!  I had some further thoughts on this and the implications for the spread of open science.  Instead of re-writing them, I'll just quote my comments on the FriendFeed thread:

I think big factors are Andy's commitment to open science and his new PI's commitment to making an impact in science and medicine.  I met Hugh Smyth a few times when he was at UNM and only detected awesomeness, both in his research and in his mentoring and concern for students.  Openness is probably going to be more challenging for them, though.  One reason is their research is much more applied and medical, and thus IP plays a major role.  The field is probably a lot more competitive.  And their lab is much more successful with funding.  As Nielsen and others have pointed out, the current reward system stacks the cards against openness.  So they will have to be careful.  But I think they're clever enough to figure out how to do it, and their success will pave a lot of roads for future openness.  I've been thinking about it pseudo-mathematically and I think the fact that they're even willing to try is a success.  I've had two PhD students graduate so far.  One is likely in industry for a long time and unlikely to be open for a long time if ever.  The other, Andy, is now at least partially doing open science.  The subsequent students in our lab (Anthony, Alex, Nadia, Pranav) are still performing open science.  A former intern, Diego Ramallo Pardo is in grad school at Stanford and has a passion for openness, but not able to be open yet.  Dozens of undergraduate lab students have performed open notebook science in my lab course, and there have been a few instances of continuing ONS after the course (most do not continue in research careers).  So, at first glance it appears that there isn't a high rate of spread of openness from our research and teaching labs.  But it occurs to me that it doesn't matter.  If we were to model openness as an infection, it's a powerful one.  I think it's even a latent infection in almost all scientists.  Participating in openness awakens the infection for life and it sheds constantly.  The immune reaction is our current system of practicing and rewarding science and it's quite powerful.  So it wins in a lot of cases.  Nevertheless, openness is slowly winning more often and the immune system is not going to adapt to get stronger.  On the contrary, the immune system is going to take major hits in the coming years.  Funding agencies are going to change rules.  Tenure and Promotion and hiring committees are going to add members who value openness.  Closed-access publishing for profit is going to topple precipitously.  And at that point, openness will spread and emerge naturally and quickly.  It seems plain as day to me.  Now, one of you all can translate that into epidemiological mathematics and fiddle with some exponents.

Friday, October 28, 2011

Open Access Week event at U. Arizona: Reproducibility, Open Data

Earlier this week I was lucky to participate in the Open Access Week event at the University of Arizona: The Future of Data: Open Access and Reproducibility.  The event was hosted by Chris Kollen and Dan Lee of Arizona University Libraries.  I am very grateful for the invite and the opportunity to meet them, some active member of the audience, and the other speakers, Victoria Stodden and Eric Kansa.

Victoria Stodden gave an excellent talk, framed around the computational sciences, and with the major point: Instead of promoting "open data," we should promote "reproducibility" in science.  She argued, very convincingly, that good science requires reproducibility and thus scientists should be easily convinced that we need very high standards for reproducible results.  For computational research, the only way to ensure reproducibility is to publish much more open data and open code than is normally done now.  If your result is computational, how can anyone hope to replicate and build upon your results if you haven't provided the source code and the data sets?  They can't, but publications without code and data are by far the most common these days.  It's a failure of science that is probably caused by many factors.  One that comes to mind is that computational scientists have been forced to fit their "publications" into standard peer-reviewed articles, where the system is not set up to accept and / or host source code and data.  (As an aside, this is clearly a routine failure of peer review, as referees obviously are not ensuring reproducibility of the research, which should be a primary criterion for publication.)  Scientists understand that reproducibility is an essential element of research.  For example, two years in a row, my undergraduate physics majors identified reproducibility as the most important element of good science (see brainstorming 2010).  Since scientists understand this, then they will naturally practice open publishing of data, code, methods when they realize that reproducibility is missing without those elements.  As Victoria argued, demanding "open data" leads to confusion and resistance and ultimately probably lack of compliance.  In contrast, demanding "reproducible research" is already a cultural norm and it naturally leads to open data and open code of the most helpful variety for reproducibility.  Victoria's slides can be found here.

The notes for my presentation can be found on linked mindmaps, starting here.  (Click on the tiny right arrows to navigate.)  My notes are probably not too meaningful if you weren't at the symposium.  In contrast to Victoria's high-level talk about policies that could make a major impact, I told a few stories about open data and open notebook science in our own teaching and research labs, and the successful impact we've had already.  I think (hope) it provided concrete example of the benefits of open science.  On the one hand, I showed that open science, especially open notebook science strongly promotes reproducibility.  This has been seen best in the undergraduate physics lab that I teach.  Students read the notebooks of other students from prior weeks and prior years.  They build upon these previous results, which allows them to get the experiment working much quicker and have more time to explore new aspects of the experiment, or to develop new data analysis methods.  They are doing real science!  I showed an example of an excellent primary notebook from Alex Andrego and Anastasia Ierides.  However, I think I also showed that open data and open science make an impact beyond just reproducibility.  This impact is in reuse and repurpose of data. I told two stories where theory and research groups already have been able to use data we publicly shared on youtube.  One group has already used our data in a theory preprint on the arXiv.  Both groups expressed delight and gratitude that our data was freely availalbe.  There are two important features of these stories.  First, both groups used our data for a purpose that we had not (and probably would not have) imagined!  Clearly the impact of our data was multiplied by being public.  Secondly, we did the easiest and simplest sharing method we could find: youtube, yet we still made an impact.  We are currently working with Rob Olendorf, a data curation librarian at UNM to vastly improve our sharing.  This will include permanent citation links, vastly improved metadata (at least 10x more than the data itself), hosting by the institutional repository (much safer than our lab server), and links to other data sets.  Reason would have it that if we could make an impact with the imperfect system we tried first, then the impact will be much higher with the data shared via Rob and the institutional repository.

The final talk was by Eric Kansa, who described the amazing work of him and his colleagues on Open Context, a platform for sharing and linking archaeological data.  His notes from the event can be found here.  And his slides are available also: A More Open Future for the Past.  Despite being far from the field of archaeology, it was easy for me to see the vast impact that Eric and his colleagues are making via the open context project.  A large amount of time, sweat, and money are expended collecting archaeological data.  Without opening these data and curating and linking these data, the potential impact is severely limited.  The Open Context team has developed a method for collecting these data, archiving them, and linking them to other data sets.  The method is very effective, and importantly requires far less work than required to collect the data in the first place.  This seemed clearly, to me, a case of the huge power of data reuse and repurpose. In contrast to computational science, the power of data reuse seemed to trump the need for open data for reproducibility.  This is not surprising, given how different the two fields are.  But it was an interesting and somewhat confusing contrast for me between the needs for open data in computational research versus archaeology.

There were several engaged audience members.  One of them was Nirav Merchant, with the iPlant Collaborative.  Victoria and I were highly impressed by the computational platform that iPlant has developed already, only three years into the NSF cyberinfrastructure project.  I was simply amazed and I couldn't do it justice describing it.  The ability to ensure reproducibility of computational research with the iPlant platform is vast.  One example is how easy it is to save an image of a virtual machine and then share this image with other users.  They demonstrated this for us and it took only a few clicks and less than a minute.  I highly recommend reading more about iPlant at their site linked above.  The iPlant team that we met was energized, engaged, and collectively brilliant.  I'd love to know how they assembled their team as they've clearly done an excellent job.  I intend to keep in contact with the iPlant folks and am even hoping that I could introduce the computational platform to my Junior Lab students this year.  I think the exposure to these state of the art and "open" tools will be invaluable for their future research.

Overall, the one-day Open Access Week event was highly successful for me.  I met some amazing people and gained a lot of clarity in my thinking about the imperative for much more openness and sharing in science. Incidentally, maybe not coincidentally, during my flights I was able to read Michael Nielsen's fantastic new book on the untapped potential of connected, open science: Reinventing Discovery.  Despite having met Michael and having heard him speak a few times, I still found the book riveting and I learned a lot.  I absolutely recommend the book to anyone interested in the practice of science!

Friday, March 4, 2011

I am maximally-skeptical that there currently exists any evidence that drinking deuterium-depleted water has health benefits or will cure disease.

Because of our lab's interest in the biophysical effects of heavy water--both heavy-hydrogen water, D2O, and heavy-oxygen water, H2O18--I received a very friendly email inquiry today.  The person suffers from a health problem and currently hopes that drinking deuterium-depleted water will help with that condition.

As a scientist and a health consumer, I am maximally-skeptical of any medical claims related to drinking deuterium-depleted water.  This is despite that fact that I think there's a good chance that cells may behave differently if deprived of deuterium, which exists in all natural water sources.  The reasoning for my skepticism is very straightforward.  There is a dearth of any published scientific or medical research utilizing deuterium-depleted water.  As I will note below, there are less than a dozen research papers on the topic.  So we really don't know.  There is almost no evidence.  We don't know whether drinking large quantities of deuterium-depleted water will be helpful or harmful or negligible.

There is much more evidence, though, that the quantity of water that would need to be consumed is quite large.  Because deuterium is natually-occurring, there's a lot of it in your body!  It would take a long time of drinking lots of D-depleted water to have a systemic effect.  My interpretation of the existing evidence is that by far the most likely outcome of this therapy is that it will generate profit for whomever is selling the D-depleted therapeutic water.

Because I think it's a shame that a Google search for "deuterium-depleted water" is overrun by claims of cures for horrible diseases, I asked the person who wrote me if I could send my response on my blog instead of privately.  So that perhaps our discussion could benefit more people.  The person kindly agreed and so I will post his email:

Dear Steve,
I enjoyed reading your blogs and noted that you work with D2O.
I have a medical condition that I want to treat with alternative methods - one of them is drinking "light" water.
Do you know, or can you suggest any resources for the following:
1. how to make "light' water, with D2O concentration of below 50ppm 2. who does D2O concentration testing in the us for water samples 3. who makes light water (for sale) 4. any scholarly literature on this topic...
Any info will be much appreciated and shared with fellow friends who are in need.
Thank you very very much!!

Here is the reply I would have sent, but instead post publicly:

Dear ___, 
Thank you for your kind message.  I am sorry to hear of your medical condition.  I am not an expert on the medical effects of deuterium-depleted water.  In fact, I am not aware of any medical experts on this topic.  As a scientist, I am maximally-skeptical of any claims of currently-known medical benefits of drinking deuterium-depleted water.  I'm not saying it will help or hurt you, I'm saying that I don't think anyone is close to knowing whether it will be helpful, harmful, or negligible.  There is almost no published, rigorous research on the subject (your question #4), and thus any claims are probably speculation.  I would suggest talking to a medical doctor, which I'd guess you've done plenty of, since they know almost infinitely more about the human body than I do.  However, I would think that any medical doctor, or indeed any living person, would merely have to guess, because I do not see any experimental evidence beyond just less than a dozen published reports which have yet to be challenged or supported. 
Below I will put responses to your specific questions, and I wish you the best, 
Sincerely,
Steve 
1. I don't know of an efficient method for producing mildly-deuterium-depleted water.  The deuterium-depleted water we use in our research is much more depleted.  We obtain it from Sigma, a chemical supply company, and it is roughly $100 per 100 milliliters (a few ounces).  As you may know, you would probably need to drink a lot of water over many days to appreciably deplete deuterium from your body.  This would surely be expensive.  And like I said above, as far as I can ascertain, it's unknown whether it would be helpful, harmful, or negligible.
2. I don't know who does D2O testing.  I'd be skeptical of anyone offering these services related to this medical purpose.  Incidentally, deuterium-rich water is inexpensive.  You could easily mix D-rich water with regular water and see if the purported D2O-testing company is able to correctly discern the difference. 
3. We so far have only purchased from Sigma.  See for example product #195294.
4. I have read two scholarly papers on the subject, both from a research group out of Hungary.  I found both papers very interesting, but I also am highly skeptical of the interpretation of their results.  A good place to find scholarly papers related to biology or medicine is on Pub Med.  This link will hopefully take you to a search for articles related to deuterium-depleted water.  I can only see one that is freely available.  Google Scholar is another place to search, but it will not be limited to biological articles. 
I actually find this topic fascinating, as far as whether life has evolved a beneficial use for naturally-occurring deuterium.  We have a side project in our lab to see whether we can notice any effects on tobacco seed growth.  We're using tobacco seeds because they are tiny, so we don't need much water to see an effect.  We got this idea from Gilbert N. Lewis, who did the initial studies in the 1930's that showed that too much deuterium affects life. One of the reasons I find this side project on deuterium-depletion so fascinating is that I see it as an open mystery.  That correlates well with my skepticism of claims related to therapeutic effects of drinking light water.

Below, I will embed a comment thread from FriendFeed, and also there are potential comments on the blog itself.  I expect them to be a mix of helpful and derisive...hopefully more of the helpful type!

Saturday, February 5, 2011

An open data success story

Over the winter break, Andy Maloney and our lab enjoyed an open data success story.  Andy shares his data publicly with a CC0 / public domain license.  Some scientists ran across the data, I think by Google searching and contacted us to ask if they could use our data to support their research.  Since it is CC0, they didn't have to ask, but like most scientists, they were courteous and did contact us.  I shared this story at the ScienceOnline2011 "Data Discoverability: Institutional Support Strategies" session and I think people liked the story.  Jean-Claude Bradley mentioned it in his blog summary of the conference, and Lucy Power saw this and contacted me for more details.  Lucy's is studying e-Research for her Ph.D. dissertation topic.  I sent her a reply, and instead of rewording it, I will just past it below.  I can answer questions on the FriendFeed thread.  Yay Open Data!


Hi Lucy – I definitely should write up a blog post about it and I will try to do that soon.  I think it’s a great little success story for open data and data reuse.  In a nutshell (and I can answer questions): Some people found Andy’s microtubule gliding assay data on youtube and emailed us to say it was very interesting to their theoretical work and could they use our data in a pre-print.  We replied “of course!” “woo hoo!” and we told them that it’s all public domain data so they are free to do whatever.  As a courtesy, we said we’d like a shout-out.  They went further and offered co-authorship, but Andy and I decided an acknowledgment was more appropriate at this time.  Andy suggested they acknowledge open notebook science, etc. and they did in their pre-print.  You can find the pre-print here: http://arxiv.org/PS_cache/arxiv/pdf/1101/1101.2225v1.pdf see Figure 3A for Andy’s data and the acknowledgments section.
 I think it’s a great success story because (A) they never would have known about our data if it weren’t open.  It didn’t necessarily have to have an open license, but it needed to be discoverable.  (B) we never would have thought to use our data for this purpose.  So obviously value was created via openness. OK, I’ll try to write up the story in a blog or something soon!  (Maybe I should just post the above and not worry about wording it better? J )
 --Steve

FriendFeed Thread:

Tuesday, September 7, 2010

Thank you, Addgene, for the award!

A big "Thank You!" to Addgene for giving Andy Maloney and our lab a "Resource Sharing Award!"  The award is a $5,000 donation to our lab that we can use to further our kinesin research.  Very generous and very helpful to our lab.  Big props also to Andy for applying for the award with no help from me!  One more piece of evidence that the students in the lab are much better at grant writing than I am :) 

The award was given to Andy and our lab for our commitment to open science.  This includes open notebook science, open data, sharing protocols, designs, etc.  Andy has been a very impressive open scientist.  It's just a guess, but I'd say so far, probably his biggest impact has been with the very detailed "do it yourself" biology projects he's contributed.  He's absolutely amazing with designing solutions from off-the-shelf components, and equally amazing with using Google Sketchup and photographs to describe the designs to the public. A good example is his microscope objective heater, which was somewhere around $500 and is working very well for our gliding motility assays.

There are now many labs around the world deserving of this award, and it feels really good to receive it.  And I think it was a great contest for Addgene to sponsor.  I actually wasn't aware of Addgene before Andy told me about the contest.  So just learning about them made the effort worthwhile.  I had a great conversation with Melanie Herscovitch on the phone a few weeks ago and she explained to me Addgene's mission and services.  Here's a picture from their website to explain what they do: (used without permission! :) )

In a nutshell, addgene is a non-profit organization dedicated to making it easier for researchers to share and obtain published plasmids.  Authors of papers submit their plasmids to Addgene (either purified DNA or transfected cells, as I understand it).  Readers who would like to obtain the plasmid contact Addgene, and Addgene provides the plasmids for just a cost-recovery fee.  This works out well for all parties.  Without Addgene, it's often a very inconvenient process.  The authors are burdened with keeping track of plasmids that may have fallen out of use.  And researchers requesting the plasmids often face a long delay in obtaining them.  I think Addgene is a wonderful service and I look forward to working with them as we create and publish our own plasmids in the coming years.  I also got the feeling from talking with Melanie that Addgene is a really great place to work.  I don't know whether or how often they're hiring, but you can take a look here for current job openings.

Thanks again, Addgene!  If you're reading this, it'd be great to post a little thank you comment (on friendfeed or the blog) or congratulations to Andy!

FriendFeed Thread:

Sunday, February 21, 2010

Science Commons Symposium Pacific Northwest 2010, quick summary #scspn

Today I was lucky to attend the amazing Science Commons Symposium.  There were back-to-back fascinating presentations by Cameron Neylon, Jean-Claude Bradley, Antony Williams, Peter Murray-Rust, Heather Joseph, Stephen Friend, Peter Binfield, and John Wilbanks.  It was wonderful to meet in person a few people that I either did not know before, or whom I'd only know online previously, including Lisa Green (who went out of her way to invite me to attend this conference, thank you!), Heather Piwowar, Anali Perry, and Brian Westra.  It was also a great pleasure to meet again people whom (with the exception of Cameron) I'd only met in person a few weeks ago at the ScienceOnline2010 conference: Jean-Claude, Cameron, Hope Leman (another tireless organizer of the conference who graciously invited me to attend), Bill Hooker, Pete Binfield, and Antony Williams.  (My apologies if I missed out on name-dropping anyone, it hurts me more than you!)

I'm off to the Biophysical Society meeting tomorrow morning, and a bit out of steam, so I'm going to cop-out a bit and just embed a mindmap of my notes from the meeting.  Before doing that, I'll post a few action items from the meeting, and maybe later I'll come back and link to specific friendfeed or other threads for the items:
  • Antony and I made some progress discussing our athletic challenge to raise money for asthma research or other charity.  I think it's promising we can make it "generative," and successful.
  • Pete Binfield used Heather Piwowar's PLoS ONE paper as an example.  I want to read it and then rate it.
  • Improve our lab's Open Notebook Science.  This is ill-defined, but there are many steps we can begin taking immediately to work towards a system that works for us as well as it does for Jean-Claude Bradley and his students / collaborators.
OK, Here's the mindmap.  I started doing it as an example for Heather, Hope, Bill and others from a discussion at dinner.  Too tired to convert it into regular text now...would be interested to know if it's useful at all to you!


FriendFeed comment thread:

Thursday, January 21, 2010

ScienceOnline2010 -- Top N things I learned and was inspired to do at #scio10

I had a blast at ScienceOnline2010 last weekend!  Thank you Anton, Bora and others who spent so much energy organizing it!  Approximately 250 people attended and it was a very diverse crowd of scientists, science writers, publishers, librarians, science outreach specialists, high school teachers, even high school students.  Much has already been blogged about the conference, including many "Top N" lists.  You can find a list of them here.  My favorite so far is Jonathan Eisen's "Enough w/ the good: here are the top10 problems w/ the #scio10 meeting."  Hilarious!  Despite everything having already been done, I hereby present my belated list of things I learned, things I did, and things I've been inspired to do:

1.  I had a ton of fun interacting with a bunch of e-friends, old and new.
I attended as a scientist with an interest in all of the other areas.  It was definitely a new feeling to be in a session and have the speaker ask, "how many of you are scientists?" with the answer being a very small fraction of the participants in the room.  It was also a very new and thrilling experience to finally meet in person many people I've known only virtually for the past year.  I think Cameron Neylon was the only person I'd met in person previously.  Despite that fact, it was incredibly easy to have conversations during lunch, between sessions, and of course at the bar.  I uniformly enjoyed these people even more in person.  An incomplete list of the people I had the pleasure of chatting with include: Bill Hooker, Jean-Claude Bradley, Pawel Szczesny, Walter JessenChristina Pikas, Hope Lehman, Peter Binfield...  Plus I found some completely new friends at the conference, including Antony Williams, Greta Munger, Dorothea Salo, Andy Farke, Natalie Villalobos, Michael Habib, Annie Crawley, ...  Now, those previous two sentences could be perceived as egregious name dropping, which I am guilty of, simply because it's quite an amazing list of people, none of whom I knew before I became active in open science.  Clicking through those names and reading what they're saying, and you'll realize why I feel so lucky to have met them!
(Note, I forgot: Fabiana Kubke, ...)

2.  Antony Williams and I made some kind of running challenge.
I knew of Antony Williams from his ChemSpider fame.  I also had recently read one of his personal blog entries, about running 1000 miles in a year, along with his ugly and unfortunate calf injury.  I'm not sure I knew these were the same person, though.  Nevertheless, he walked passed me in the hotel bar, and I accosted him to inquire about his calf injury.  I'm pretty sure he didn't know me at all, but luckily we had on slick name badges and I was surrounded by credible people.  It'd be a big challenge for me to recount the conversation (I swear I remember it perfectly, I just don't feel like writing it down).  Let's just say that I was happy to learn that he'll be back running again within a week or so and that he has a goal of raising money to fight asthma.  I have had an idea mulling in my head that I could raise money and get motivated to get better at running by setting a race time goal.  I thought this was a perfect match with Antony's goal, so I quickly challenged him to a running competition.  He quickly agreed (fearlessly) and I tweeted/friendfeeded it to lock in the deal.  Over the next week, I'll see if I can clarify the challenge and I'll post updates to that thread.  I'm thinking I'll setup a Google spreadsheet for me and others to place their pledges and monitor the progress.  Antony already knows about Nike+ technology, and I'm looking forward to doing something like that too.  It'd be a good way to try out new things in open data, and open notebook science, actually.  Suffice to say that I'm going to get better at running, lose a lot of weight, and hopefully we'll raise some money too!

3.  I and KochLab are going to get better at doing Open Notebook Science and sharing data & software this year
I learned a lot at the conference about tools that exist for carrying out open notebook science and sharing data, methods, software, etc.  I still have a lot to learn, and indeed many tools still need to be developed.  But I know that our lab can make improvements this year.  Here's some concrete things that we'll do:
  • I've sent an email to Amy Jackson, Digital Initiatives Librarian at UNM, requesting a meeting.  I had briefly spoken to her via email in October and now I'm fully energized to have a meeting with her and see what kinds of first steps we can take towards building a partnership between the library and our lab in terms of sharing data and conducting open science.  I'll try to leave updates on this FriendFeed thread.
  • Get better at sharing software.  Currently we use LabVIEW, which is a wonderful programming environment.  One of the major benefits is that it's a graphical, data flow language.  This makes the code a 2-dimensional diagram...so in my opinion, it's exceptionally easy to read other people's code.  Unfortunately, it's a proprietary and expensive coding environment.  You are allowed to compile .exe and .dll files for others to use freely.  But that's not open source.  So, there are two routes to go: (1) We could compile virtual machines (VM) and send those to people for exploring our code.  For example, referees of papers we submit.  This was an idea from Deepak Singh at the meeting. Licensing is an issue here, and what I'd like to do is find someone at National Instruments (creators of LabVIEW) and discuss what can be done to serve our open source needs.  (2)  Learn a text-based, freely available language.  I think this would be valuable for our students anyway, in terms of building their resumes.  While at the meeting, I thought Ruby was a good idea, but now not so sure.  I've posted a FriendFeed message about this, and have received all kinds of very valuable advice.
  • Adopt techniques to make it easier to capture our workflow in the lab.  OpenWetWare has innovations coming up soon and we'll certainly jump on those.  Cameron also hinted at something revolutionary coming up this year, but said he'd have to kill Bill if he told him what it is.  It wasn't clear that he'd have had to have killed others in the room, so I was disappointed he didn't say what it was.  (Just kidding, Bill!)   But the fact is, I don't think our current tools are nearly sufficient.  I'll put in more effort to make positive steps here, but I'm not sure exactly what yet.
4.  Improve the state of publishing, one article at a time: Try out some ideas via PLoS ONE.
I was delighted to have the opportunity to talk with Peter Binfield at the bar and discuss publishing ideas with him.  I have a lot of ideas that I'd like to try out, and it hasn't yet been proven that all of them are lousy.  I ran many of them past Peter and his general response was, PLoS ONE would love that -- it's just that whatever scientific community you're in may not respect it.  Very true.  However, I do think many in my community would respect anything that showed how publishing could be better by being different.  And I now realize that PLoS ONE is a very good platform for trying out a few things.
  • At lunch with Bill and Pawel, they mentioned an idea that others had been talking about.  Unfortunately, I don't know to whom to attribute the idea. (Note: Bill and Pawel give attribution to Fabiana Kubke, which sounds right to me now.)  The idea is that publishing one very good figure would be a good idea.  I like this idea enough to think about trying it out.  It is related to an idea Larry Herskowitz and I were discussing a few days before the conference:  Can we publish a paper without an introduction?  Or an introduction that just says, "see such-and-such other paper for introduction?"  So much time is wasted rewriting introductions (in my opinion), especially when trying to avoid copying prior written work.  Is this necessary?  Publishing a single figure takes this even further.  I have some interesting data from grad school and postdoc that I have not been able to publish.  It pains me that it's just sitting around, useless.  Do you think it's worth using this data to test out the "publish one figure" method?   If I make progress on this, I'll post updates on this FriendFeed thread(Note added: Cameron Neylon reminded me that BMC Research Notes may be a better venue than PLoS ONE for this.)
  • For maybe ten years now, I've thought that the value of anonymous peer review is overstated.  During that time, I've heard other people, much more eloquent than me also express this opinion.  Just briefly, I think fully-attributed, non-anonymous peer review would solve many problems that exist with today's science, and I discussed this a bit at the meeting.  Two of these problems are: (a) good referee work is difficult, and good referees are not credited for the work and the original ideas, and (b) a whole lot of incomplete and sloppy work is submitted and much of it is published due to ineffective referees / editors.  The solution I like is for every aspect of the peer-review process to be published, including the original manuscript submitted, all subsequent revisions, and all communication between authors, editor, and referees.  Clearly this solves problem (a).  As for problem (b), people often say, "but people aren't going to say negative things if their name is on it!"  First of all, that's not necessarily true.  Secondly, referees have the option to decline without comment.  If the editor cannot find someone willing to slam the paper, then it's just returned to the authors.  Voila!  One less crappy paper published.  The arguments get more complicated, especially when considering that grant review is anonymous, providing a lot of opportunity for underhanded retaliation.  I can't mathematically prove that it's a good or bad idea, so I'd like to try it out and see what happens.  It occurs to me that I could submit a manuscript to PLoS ONE and request that the Academic Editor implement this idea.  Why not?  Should be OK as far as I understand the rules.  I don't expect to submit garbage, so it won't be a complete experiment.  But something may be learned, and at least all the referees will get credit.  I may update progress on this idea on this FriendFeed thread.
OK, That's enough for now.  I can always add more later, especially since I aptly named this post.  For example, I may talk about the rebranding of our blogs that Walter Jessen recommended.  It's a good and valid suggestion, but I'm not adding it yet, because I'm not sure how soon I'll be able to think about that :)


Thursday, December 10, 2009

A PLoS ONE Success Story--Taxol Crystals Masquerading as Microtubules


ResearchBlogging.org
Andy Maloney, a Ph.D. student in our lab, recently read and summarized a very interesting paper in his open lab notebook. The paper, "Taxol Crystals Can Masquerade as Stabilized Microtubules," was published in PLoS ONE in January of 2008 by Margit Foss, Buck W. L. Wilcox, G. Bradley Alsop, and Dahong Zhang1. Since our lab is now heavily involved in experiments involving kinesin and microtubules, and because it addresses something that had been a mystery to us, the paper really caught my interest. I'll explain more about that below. But before doing that, I wanted to talk about something probably of more general interest: a success story for publishing in PLoS.

Andy noticed that in their methods they defined BRB80 as having 4% glycerol. Glycerol is used to promote tubulin polymerization, and I've never seen it included in the BRB80 (aka PEM) definition. It could also affect solubility of Taxol, so it's an important detail whether or not a substantial amount of glycerol was in their standard BRB80 buffer. I strongly suspected that this was just an oversight by the authors...and I could easily have assumed this and moved on. But what about future readers of the article? Was there anyway to correct that article? For most journals today, even in the year 2009, the answer would have been, "no." However, this is no ordinary journal, this is PLoS ONE! All I had to do was select the text in question, and then click to add a note. After adding my note, an icon appeared in the article, allowing any future reader to see the question.
PLoS Comment Image

I don't know whether authors are notified when their article is commented on. (If not, it would be an important feature for PLoS to add.) So, I sent an email to the corresponding author of the paper (D. Zhang) pointing out the question. In less than a day, D. Zhang wrote back saying that he'd asked M. Foss to look into the issue. And then again in less than a day, Margit wrote me back to say that she'd looked at the original lab notes and indeed they'd made a bit of a typo in how they described BRB80 in their report. She added a very clear response to my note. She also went out of her way to point me to two subsequent papers that have extended their taxol microcrystal research2,3. These authors deserve a lot of praise for responding to this question so quickly! A few months ago received a similarly rapid response from authors of another PLoS article...only two data points, but I wonder if PLoS authors are indeed more likely to respond quickly to questions from readers?

Now, why am I so happy and why do I think this is a success story for PLoS? It's because now, for the rest of time, when readers of this excellent paper do look into the methods, they will be able to see the corrected definition of the buffer used. Given how many times I've been burned by incomplete or incorrect methods, I do believe this will save substantial amount of time for at least a couple people down the road. (Will the PDF version of the article ever incorporate this note? As it stands now, I don't think it does...it would be very valuable if technology could be worked out to include links to these comments in future PDF downloads.) One more thing: I just noticed that Margit Foss today also posted a new comment on her article. She links to the two papers she'd told me about in her email, as "Relevant references on Taxol crystals." This is a great service to readers, especially since the newer reports2,3 support a different mechanism for Taxol microcrystal / fluorescent tubulin binding. In summary, many thanks to PLoS for this wonderful journal and to these authors for their dedication to excellent science!

Now, if you're still reading, I'd like to also comment on the very interesting science in their report. Taxol (generic name is paclitaxel, I think) is a drug used in cancer chemotherapy. It's proposed mechanism of action is to inhibit mitosis by stabilizing microtubules in the spindle apparatus. In vitro, Taxol dramatically reduces the rate of microtubule depolymerization. Many people, including kinesin researchers in our lab, leverage this microtubule-stabilizing effect by adding Taxol to microtubule-containing solutions. What I learned from the Foss et al. paper is that the concentration of Taxol typically used in microtubule gliding assays (10-20 micromolar) is far above the solubility limit of Taxol (somewhere around 0.8 micromolar in aqueous solutions). Furthermore, they show that Taxol forms microcrystals above this solubility limit (even at 0.92 micromolar) and that often these microcrystals form a striking resemblence to microtubule bundles and asters! DIC images of these microcrystals (formed in absence of tublin) are shown in these images from Foss et al.1:

(scale bar 10 microns)


The final piece of crucial information provided by this article is: these Taxol microcrystals rapidly bind fluorescently-labeled tubulin! (Later reports indicate that it's the fluorophore, not the tubulin that is binding to Taxol2,3.) This means that many kinesin researchers (including me) likely have Taxol microcrystals in their samples, and because they become coated with fluorescent tubulin, there is a huge risk of misidentifying these structures as microtubule structures. Indeed, here is a recent fluorescence microscopy video that Andy took of something that at the time was a mystery but which we now know is likely a Taxol microcrystal decorated with rhodamine-labeled tubulin!

Likely Taxol microcrystal in kinesin / microtubule gliding motility assay (using rhodamine-labeled tubulin). Andy Maloney data.

In my past, I've also often seen these structures which I attributed to "clumpy" or "weird" microtubule structures. For example, I often noticed very bright, thick, and stick-like structures that I called "microtubule logs." It never occurred to me that they were Taxol crystals! (Also I remember that these structures were much less prone to photobleaching. I wonder if that's because (a) there are buried fluorophores inside the crystals, protected from oxygen, or (b) even on the surface of the crystals, Taxol somehow protects fluorophores from photobleaching?)

Foss et al., go further and speculate on whether this has important implications in vivo (i.e. in cancer chemotherapy). I can't really comment on that, but it's interesting to think about. What's most important for us is that we now know we have a problem with our buffers (too much Taxol!) and we may be able to fix it. The concentration of tubulin that we typically use is about 0.4 micromolar of tubulin dimers. Thus, for a 1:1 ratio of Taxol to tubulin dimers, we'd need 0.4 micromolar starting concentration of Taxol, which is below the solubility limit. There's at least two things I don't know: (a) What is the binding affinity of Taxol for microtubules? and (b) Do we need a 1:1 ratio to get significant stabilization? If the answer to (a) is something like a few nanomolar, then we may be OK with something around 0.5 micromolar (500 nanomolar) Taxol. If not, then we may have to hope the answer to (b) is "no."

A quick search just now yielded a paper from 1994 that says the binding constant for taxol to microtubules in 10 nM. That'd be good, except that they also seem to say that they only get stabilizing effects when the concentration is in the micromolar range4. Dang! Well, it shouldn't be too hard to try out 500 nM Taxol and to see whether MTs are reasonably stable. It's possible our MTs may be more stable than those used in the Caplow et al. study. It's also possible that the Taxol microcrystals are not affecting the kinesin activity in our system, and that we can do our studies at high Taxol concentration. Even if so, it's great to know about this issue so we can keep on the lookout for Taxol problems.

References

1. Foss M, Wilcox BWL, Alsop GB, Zhang D (2008) Taxol Crystals Can Masquerade as Stabilized Microtubules. PLoS ONE 3(1):e1476. doi:10.1371/journal.pone.0001476

2. Castro, J. S., Deymier, P. a., Trzaskowski, B., & Bucay, J. (2009). Heterogeneous and homogeneous nucleation of Taxol crystals in aqueous solutions and gels: Effect of tubulin proteins. Colloids and surfaces. B, Biointerfaces. doi: 10.1016/j.colsurfb.2009.10.033.

3.
Castro, J. S., Trzaskowski, B., Deymier, P. a., Bucay, J., Adamowicz, L., Hoying, J. B., et al. (2009). Binding affinity of fluorochromes and fluorescent proteins to Taxol™ crystals. Materials Science and Engineering: C, 29(5), 1609-1615. doi: 10.1016/j.msec.2008.12.026

4. Caplow, M., Shanks, J., & Ruhlen, R. (1994). How taxol modulates microtubule disassembly. The Journal of biological chemistry, 269(38), 23399-402. Retrieved from http://www.ncbi.nlm.nih.gov/pubmed/7916343.

Foss M, Wilcox BW, Alsop GB, & Zhang D (2008). Taxol crystals can masquerade as stabilized microtubules. PloS one, 3 (1) PMID: 18213384

Link to FriendFeed discussion thread.

Tuesday, June 9, 2009

My first rating and commenting of a PLoS article in my own field (Scary!)

SJK 6/9/09: Here is a link to related friendfeed discussion.

I just finished reading and commenting on a PLoS One article that is near my own field of research. The article is titled, "Dissection of Kinesin's Processivity." The authors are: Sarah Adio, Johann Jaud, Bettina Ebbing, Matthias Rief, and Günther Woehlke. You can see my rating and overall comments here. (Since I'm not sure if that link will work, I'll also repost my comments below.)

Throughout the process of reading and commenting on this article, I learned a lot more about my fears and barriers to PLoS commenting. I discussed some of these in my prior post about my first PLoS rating. In contrast to my first rating, this article is smack in the middle of my field of interest (the kinesin molecular motor). I deliberately chose the most relevant PLoS article I could find. I'd estimate that my fear of placing comments was at least 10 times higher than for an article outside my field. I definitely felt like my comments were piping directly into the author's email inbox, ready to enrage them at any misunderstanding or criticism I posted. I still feel this way and am a bit worried. My worries are probably justified to some extent, since I am very new to this field. Thus, I could easily be seen as an ungrateful newcomer who hasn't paid his dues. And of course the people who wrote the article could end up anonymously reviewing my own papers and grants.

Given those worries, I came close to deciding not to post my rating. However after much reading and thinking about their results, I felt compelled to make a serious comment about error analysis supporting one of their conclusions (not their major conclusion). I was confident that my criticism was fair, and convinced myself that posting the comment was the right thing to do--perhaps I can save another reader a lot of time, or even help the authors out if they read it. I posted my criticism directly in the article, along with several typo corrections. After doing that (late last night), I realized that if / when the authors DO see my comments, they'll see a string of petty typo corrections and then this criticism, but nothing positive at all. That's a problem!!! Because of this, I decided to sleep on it, and compose an overall rating with positive comments today. I was busy most of the day, but finally tonight was able to finish my rating. In all honesty, though, without having travelled that slippery slope of commenting, I don't think I would have posted this rating tonight. I would have balked at the risk of angering the authors, sticking my neck out, and possibly being wrong. I probably would have convinced myself that these risks outweighed any meager potential gain that the world of science would get from my remarks.

I'm a bit worn out now. Hopefully in the comments here or more likely, on FriendFeed, we can talk about these things. I hope in the next couple days to expand on my review of the paper in my research blog, and to include it as my first Research Blogging attempt.

Reposting of my rating and overall comments on the article

This is what I submitted to PLoS as my rating:

Insight: 4 stars, Reliability 3 stars, Style, 4 stars.

The authors recently characterized NcKin3, which is the first known,
naturally dimeric but non-processive and plus-end motor. In this
report, they are leveraging this discovery to study chimeric constructs
between NcKin (a dimeric, processive Kinesin-1 motor in the same
organism) and NcKin3. They make two different chimeric constructs: one
with the head of NcKin and the neck of NcKin3, and the other with head
of NcKin3 and neck of NcKin. Importantly, the head included the core
motor domain AND the neck linker region.

I congratulate the
authors on a lot of very nice work that must have been very difficult!
The results they report come from an impressive array of difficult
assays spanning single-fluorophore position tracking, single-molecule
bead motility assays with optical tweezers, gliding assays, and a
variety of ensemble biochemical assays.

Study of the two
chimeric constructs, in comparison with the NcKin and NcKin3 wildtypes
allowed the authors to gain insight into which parts of the kinesin
motor are important for conferring processivity onto dimeric
constructs. (And also, which parts are important in NcKin3 for
inactivating one of the heads.) As far as I know, these are the very
first two chimeras created between these two kinesins and thus open the
door for many more investigations into how processivity is regulated in
the motor domain, neck-linker, and neck regions. The results here
indicate that many more chimeric structures and site-directed
mutagenesis studies will be necessary and valuable. Of course, that is
a lot of work, but the results here open the door for those further
studies.

For me, the most fascinating result was point (iii) on
page 4. The authors show that the Head3/Neck1 construct seems to get
stuck in a "kinetic dead end." As they say, the kinesin-1 neck appears
to confer some elements of processivity, but not all. Combined with the
missing elements (which kinesin-3 head lacks), the motor is actually a
bit more handicapped, as shown by a gradual decrease in gliding
velocity as the concentration of motors is increased.

I also had a couple questions about the paper that I noted previously (see prior article comments):

* Statistical significance of processivity measurements.

* Lack of discussion and comparison with previous Ncd/Kinesin-1 chimera results

DISCLOSURE:
Our lab (http://openwetware.org/wiki/Koch_Lab) has recently obtained
major funding to study kinesin. I do not think we have competing
interests with these authors or the work they've presented here, but I
thought it worth mentioning.

Thursday, May 28, 2009

Pondering our Kochlab graduate student compass...how about "Always Contribute?"

Last weekend, my friend Richard Yeh posted a couple essays by Paul Graham onto Facebook. I loved the essays and linked to one of them on friendfeed. Michael Nielsen, in turn pointed me to another essay by Graham that he thought I'd like, "How to Do What You Love." Michael was completely right, I loved the essay. If you have not read that essay, I command you to stop reading my blog and to go read that article! You'll get much more out of his essay than this blog.

OK, now that I have you defiantly reading my blog, intent on garnering something useful from it, to spite me, let me continue. The "Do What You Love" essay resonated with me very strongly. It reminded me of the discussion of talents in "First Break All the Rules" by Buckingham and Coffman. I think Graham and Buckingham and Coffman are talking about the same thing: that finding work you love is a key to happiness (and productivity), but that finding out what you love is a very difficult task worth working very hard on. The language of Buckingham and Coffman is to talk about finding one's "talents." I've been talking with my graduate students about this a lot for the past six months. (In fact, it's time for me to have another awkward talent-finding session with them, I do believe!) I also preach to all of my undergraduate students about the importance of finding their talents and I give them an end-of-semester assignment to think about their talents. I'm delighted to have been shown the Graham essay, because I think it is yet another way of presenting this argument to students, and a very eloquent one.

Since I loved the essay so much, I sent it to the person who gave me the First Break All the Rules book. He wrote back to me and keyed in on the "always produce" part of the Graham essay:

"Always produce" is also a heuristic for finding the work you love. If you subject yourself to that constraint, it will automatically push you away from things you think you're supposed to work on, toward things you actually like. "Always produce" will discover your life's work the way water, with the aid of gravity, finds the hole in your roof.

While reading the note, the Do What You Love essay finally clicked with another I read by Graham last weekend, "How to Make Wealth."[1] It's another fantastic essay that I feel like commanding you to read. One premise in that essay is that people in start-up companies can be 20-30 times more productive than they can in an ordinary 9 to 5 job. Thus, a small group of people can create a tremendous amount of wealth by working really hard for a few years. They can also get financially rich as a reward for their production of wealth for the world. The thing that clicked for me is that you cannot make the world a better place without producing. Most people are producing at a rate at least 20 times less than they could be producing, if they found what they loved and were able to do it all the time. I garner great optimism from this fact that we're on average so incredibly inefficient. It means most people are not even close to any absolute point of diminishing returns, and with the right kinds of changes, they could easily multiply their productivity and impact on the world by manyfold.

So, then I started thinking about our research lab and the students in our lab. I thought over things that each student has done in the past year that made me profoundly happy. As I thought over all these things, I realized they had a common theme: I was recalling instances of those students being unusually productive. Furthermore, my favorite recollections involved those where the students had shared their work on our public wiki, or our blog, or in some other fashion open to the world. This made me think that it is now very easy for me to summarize my main expectation and goal for my students: "always produce," borrowed from Paul Graham, of course. My students like to make Kochlab slogans, so I thought of "Kochlab: Produce" or "Kochlab: Always Produce," but if you pronounce the "Koch" correctly ("Cook"), then it has the problem of making one think of produce the noun, e.g. apples and bananas. Thus, I am thinking something like "Kochlab: Contribute" or "Kochlab: Always Contribute."

In some sense, "always contribute" means the same thing as "always produce." The point of the producing, in regards to Graham's essays is that you're creating wealth, and therefore contributing something to society. However, I like "contribute" much better, because it has much more clarity in the science world. "Contribute" automatically points the way towards open science (aka Science 2.0). Whereas, production in the traditional scientific world ("closed science") can be done with a very limited amount of contribution.

I have mulled it over for a couple days now, and I think I really like this as the main piece of advice and constant guidance to give to our students: "always contribute." Does this work? Let me try it out in a few ways:

1. By the time the students get their Ph.D.s, I want them to have learned a tremendous amount about what their talents are. I want them to clearly see what the next step in their career should be in order to leverage those talents and help them be successful and happy. A compass of "always contribute" will lead the students towards finding ways of being productive instead of spinning their wheels. These activities will be the means by which the students and I discover what their talents are. This is the point of Graham's "always produce" advice. Check.

2. By the time the students get their Ph.D.s, I want them to have a strong and large professional network of people that know them and the work they have done. "Always contribute" tells them that Open Notebook Science is a good thing to do. Sharing code, design drawings, personal summaries of research papers, tips and tricks on protocols -- these are all ways to contribute. In our limited experience in our lab, we have received validation after validation after validation that open contributions get attention. We can see this vaguely via page views or Google search rank or quite vividly via positive feedback from people that we admire and people whom we've helped. Combined with traditional publishing (also a contribution) and attending scientific meetings (contributing), I think "always contribute" will make building a powerful professional network almost automatic. Check.

3. I want our lab to produce innovative, exciting, and high-impact scientific results. Will "always contribute" point us in the right direction for this goal? Does it point in any direction? I need to think about this one some more. I feel like it must point in the right direction--for example, innovations are contributions. But there's some risk that focusing on contributing could lead towards a lack in overall production. Basically I am thinking of the standard arguments against open science -- increased likelihood of scooping, which in turn reduces chances of funding and publishing. I fundamentally believe that those arguments are strong enough to tip the balance, but I don't think they've been proven yet. Another example of how "always contribute" may be counter to our lab's scientific productivity: some students may discover that they are wickedly talented at contributing in ways that do not advance their research projects. That's a great thing to discover! An example that hasn't happened in our lab yet would be for a student to discover that they're fantastically talented at writing popular science articles and want to do so at the expense of doing any research. I want students to discover something like this. It fits perfectly with items 1 and 2 above. However, it is clearly a problem in regards to item 3. This is not a new problem, though. My job as a research professor is to both mentor students as well as ensure production of research results. The way the system is set up those goals are not always aligned, and sometimes in conflict. It's possible to be rewarded with research grants, even by abandoning the best interests of your graduate students. Most people in the system know this and have seen the devastating results it has for too many Ph.D. students. I am absolutely against doing that and despise many people who have chosen that route. On the other hand, I don't have a good idea about what do do if "always contribute" turns into "I can't do my research." That's definitely going to happen eventually. In many cases, it will be possible for the student to discover their true calling in life, but then re-focus on making the research contributions necessary to finish the Ph.D. that they've invested so much time in. Will there come a time when the student should rightly choose abandoning the Ph.D.? Ugh, this is a tough one: Item #3 gets an: almost check / need more thinking.

I've painted myself into a corner now. If I were Paul Graham, I would figure out a way to backtrack. But I'm not, and I don't really know how to end this blog post, so I think I'm going to end it by linking to another Graham essay about writing essays. This one was linked to me by Kartik Agaram on friendfeed. It's an essay that explains why high school and college writing assignments sucked so badly. If you hated those assignments but never quite knew why, you'll love this story. Plus, you'll feel vindicated and it will give you one more reason to trust your gut in the future. For example, if your gut were telling you that "always contribute" is a fantastic compass to present your graduate student mentees.

Footnote:
[1] I'm taking a bunch of liberty here with my own story. It wasn't until I started writing this blog that I realized the two essays had clicked. But I think subconsciously this is what was happening. Also, I probably had the Gin, Television, and Social Surplus essay by Clay Shirky in my head, as Joelle Nebbe had linked to it recently.

Tuesday, May 26, 2009

My first PLoS comment: High rating of an article on TSLP being the cytokine link between eczema and asthma

5/27/2009 SJK Note: After I wrote this, Bora Zivkovic sent me links to the PLoS community blog where he talks about commenting and rating PLoS articles. Both are very much worth reading! Bora is the Online Discussion Expert for PLoS.


Recently, William Gunn Mr. Gunn composed an excellent article discussing online identity and the making of public comments in scientific circles. Without immediately spiraling into a stream of ridiculous conversation, I can't really comment on his post, or the ensuing friendfeed thread. Suffice to say that Mr. Gunn and others on friendfeed inspired me to be a lot bolder in commenting on PLoS articles.

So, tonight I made my first comment on a PLoS article. Previously, I had viewed commenting on the actual article site as a very formal procedure that required attaining the highest level of understanding of the article before submitting a comment. Essentially, I was viewing commenting on an online article the same way I viewed submitting an official comment to an article published in Science or Nature (or other journals). Published comments in those journals are almost always refutations of the article that seemingly without fail lead to concomitantly published rebuttals by the original article authors. Thus, the culture of commenting on articles is fraught with nastiness and putting one's scientific reputation on the line. This could be the reason that so far "official" online commenting on peer-reviewed articles has been very limited, whereas "unofficial" or off-site commenting has been more common. By "unofficial," I am loosely referring to comments made anywhere that is at least one link removed from the actual published article site. For example, an external blog, friendfeed discussion, or notes left on article managing services such as citeulike.

It occurred to me while laughing and crying my way through the recent friendfeed discussions (OK, fine, here's a link to perpetuate the madness) that this culture may be relatively easy to change. (Aside from any questions of whether it's necessary to change.) In my opinion, PLoS has already made one innovation that vastly increases the odds of a user making a public "comment." They have separated the article ratings into three categories: Insight, Reliability, and Style. From my personal experience, that opens the door almost all the way in terms of inviting some kind of reader feedback. Rating an article on "Style" does not carry much professional risk from my viewpoint. Rating on "Insight" requires understanding of the possible impact of the article, and is thus much more weighty than the "Style" rating. However, I personally feel I can rate an article on "Insight" without assessing the quality or reliability of the methods and data. I recently did this with a PLoS ONE article I saw on single-cell sequencing of uncultured organisms. To rate an article on "Reliability," I feel requires the kind of in-depth understanding that would be required for me to send a formal letter into the editor of Science or Nature that could be published. Thus, the barrier for me to rate on "Reliability" is quite high. Especially since if I'm going to put in enough effort to feel completely justified in rating, it's likely to be less than a 5-star rating. (I guess I'm feeling like I spend more time reading articles that I disbelieve than those I do believe?)

Another reason that placing online comments does not have to be as formal and negative as with traditional published comments is that the comments are published without a delay waiting for the original authors to compose a response. This then reduces the expectation that the publishing authors must respond and therefore takes the formality down a bunch of notches in my opinion. Also, in terms of PLoS the whole mission of the journal is to make research more broadly and rapidly available--and thus I think there is an expectation that the comments should also come from a broader base of readers.

So, that is what inspired me to take the time to read a PLoS Biology article and compose my first online comment tonight. I was also inspired by the belief that we're still very early in the process of dictating the culture of online discussions of peer-reviewed research--and thus a concerted effort can make impact in what ends up happening. This inspiration was combined with the coincidence that my wife sent me an article from BabyCenter today that caught my interest because it was discussing the recent PLoS Biology article. Finally, the thing that finally tipped the balance and convinced me to take the leap and make my first PLoS comment was a healthy dose of "WTF" So I stopped worrying and took the leap. :)
 
Creative Commons License
This work is licensed under a Creative Commons Attribution-Share Alike 3.0 Unported License.