The Hidden Costs of Cloud Storage

Cloud storage is an increasingly popular way to store research data. Being able to upload and access files from any location is useful and makes transfer between computers much easier. But for all of the upsides of cloud storage, there are also a few downsides.

Data Ownership

While most of us don’t usually read terms of service agreement, it’s worth doing a little digging when it comes to your cloud reader. For example, Google Drive’s terms of service includes this little tidbit:

When you upload or otherwise submit content to our Services, you give Google (and those we work with) a worldwide license to use, host, store, reproduce, modify, create derivative works… communicate, publish, publicly perform, publicly display and distribute such content.

You retain intellectual properties rights over the content you put on Drive but Google can still do a lot of things with your content. This should make you a little worried about any research data you put on Drive.

There are some ways around this problem. One example comes from UW-Madison, which has negotiated Google Drive terms of service for faculty and students where Google has no ownership or use permissions. The other option is just to pick a cloud storage provider that won’t use your data, but even that isn’t always perfect. Dropbox, for example, doesn’t take quite the same liberties Google does with your data but does it spell out in its terms of service how it can use your personal information (name, address, log-in information, etc) or provide your files to law enforcement.

My best advice? Read the terms of service before choosing a cloud storage provider for your data.

Security

The other natural concern when giving your data to a third party is security. This is especially important when putting sensitive information or student information (covered under FERPA) in the cloud. You need to take a lot of extra precautions in the cloud if your data is sensitive.

One secure cloud storage option I’ve run across is SpiderOak. Unlike other cloud storage options, SpiderOak cannot actually read any of your data because it gets encrypted before it even arrives at the SpiderOak servers. And in this Ars Technica review , SpiderOak favorably compares with other popular cloud services like Dropbox and SugarSync.

So unless you find such a service like SpiderOak that guarantees security, the cloud is not the best place for your sensitive data.

The Limitations of the Cloud

Cloud storage can be a blessing in the laboratory, but putting your data in the cloud does not automatically mean that your data is well backed-up nor well managed. This is because your data is outside of your control when you give it to another entity. If your cloud storage provider folds or suddenly changes their terms of service (as seen in the recent Instagram debacle), you could suddenly be in a tight spot. For safety’s sake, it’s better to have other back-ups besides your cloud drive.

I’m in no way saying that you should not use cloud storage for research. Instead, you should be smart about choosing a service provider and know that service’s limitations. With a little bit of forethought, cloud storage can be a valuable asset in the laboratory instead of a potential security hole.

Posted in dataStorage, digitalFiles | 2 Comments

The Proper Pen

Have you ever wondered what the scientifically optimal writing utensil is to use in your lab notebook? No? Well, this post contains the answer to a record-keeping question you never thought to ask.

The answer comes from one of my favorite books on managing laboratory records, Writing the Laboratory Notebook by Howard Kanare. It was published in the 1980’s (making the section on electronic record keeping highly entertaining) by the ACS and thoroughly covers the how’s and why’s of keeping a proper notebook.

This book is so thorough, in fact, that it spends 6 pages (p. 11-16) on the proper type of paper and ink to use. Kanare even conducts experiments with 15 different types of pens to determine the most colorfast and solvent-fast inks. I found his experiments so interesting that I thought it worth sharing the highlights with you.

Just say no to pencils

First, I should say that pencils are right out. They’re erasable, they smudge, and they don’t copy well when you’re backing up your notebook. If you want to be sure that data hasn’t been changed or lost to illegibility, it’s better to stick with a pen.

Ink color

The choice of ink color comes down to lightfastness, since modern inks no longer contain the harsh acids that eat through paper over time–a historic problem. Kanare tested ink under both fluorescent light and sunlight and found that red inks fade most easily, blue ink fades some (the amount of fading depends on the pen type), and black inks fade the least.

Pen type

Felt-tip pens have a few things going against them from the start. Their inks are water based, making the ink more likely to bleed and less permanent. On the positive side, these porous-tip pens held up to Kanare’s solvent tests (using water, hexane, HCl, acetone, and methanol) about as well as the ballpoint pens.

The other main option, a ballpoint pen, does pretty well under Kanare’s solvent tests and the pen’s solvent-based ink makes writing more permanent. Kanare’s only warning about these pens is that the ink can coagulate or settle during long-term storage, leading to performance problems in older pens.

Kanare also brings up the option of using archival quality pens, but it’s not clear without testing if it’s worth the added expense over the long term.

And the winner is…

You can’t go wrong with a humble black ballpoint pen when writing in your lab notebook. This ink will stand up the most to fading and spills and provide good permanence, making your records readable for a long time.

Posted in labNotebooks | Leave a comment

Why Should I Share My Data?

I’m going to be talking a lot about data sharing on my blog, so it’s worth investigating why I believe sharing data is beneficial. There are plenty of reasons for and against sharing–I will highlight some of them in this post–but I believe that the overall balance comes out in favor of sharing.

Reasons for sharing

Many of the reasons for sharing are driven by the desire to make science more transparent. Data sharing helps ensure that we are conducting research properly and that our analyses are reproducible. Freely available datasets (and code!) allow others to test data for anomalies and analyses for validity (example). It is frightening to let others delve into our data to look for errors, but this ultimately makes our science better.

Another reasons for data sharing is the ability to conduct novel analyses on datasets. For example, meta-research brings together a variety of datasets, looking for connections that can’t be found in one dataset alone (example). It’s also possible that your data is useful ways you’ve never dreamed.  Freely available data lets researchers create interesting mash-ups that can lead to new science.

Lest you think that data sharing is only good for others, there is also evidence that data sharing increases article citation counts. Data sharing also gives researchers a way to get credit for traditionally unpublishable results. Just because your dataset isn’t interesting enough to publish doesn’t mean it doesn’t have value.

Finally, data sharing benefits society as a whole. Data sharing represents a return on the public’s investment in federally supported research. It also lowers the barrier of entry into research for non-scientists. Finally, data sharing without cost or barrier to access can help spread scientific ideas faster.

Reasons against sharing

One of the biggest concerns about data sharing is being scooped. If I share my data, the fear is that someone will use my data to publish my study before me. There are two points to make here. The first is that data sharing should not be expected before the data’s corresponding paper is published, preventing others from publishing your data before you. Secondly, when someone uses your research outputs (ideas or data) without proper attribution, that person is committing a transgression. I don’t think that we should avoid doing something beneficial just because some people will never follow the rules.

Another major concern about data sharing is that it hurts researchers who invest significant time and effort into their datasets. A large dataset that takes years to acquire and may be used for several papers is not easy to freely share. I don’t have a good answer here other than it is worth having a discussion on embargoing data for a short period of time.

Finally, many datasets have issues that prevent sharing, such as human subject information and health information. This is a valid concern and such data should not be shared as is. It is possible to deidentify some datasets, but I recognize that there are other datasets that just can’t be shared.

Why I believe in data sharing 

I think it’s important to remember in the context of data sharing that early scientists were secretive and did not publish their results in journal articles. This changed in 1665 with the arrival of the first scientific journal. Since then, the journal article has become the currency upon which scientific exchange is based–but the journal article is only the norm because we as scientists have made it so. If we find value in shared data, we researchers can change our norms to make data another research currency.

So why would we place research data at the same level as the journal article? We should because sharing data, with some limitations as for privacy, benefits the greatest number of people. Scientists benefit through reproducibility, novel analyses, and more citations while non-scientists benefit through better access to science and the ability to access the results that our tax dollars have paid for. Technology has enabled us to share data with unprecedented ease and, by doing so, we can dramatically further the cause of science.

This post represents the highlights of why I think data sharing is beneficial. I understand that not everyone agrees with my view and for that reason we should move toward more data sharing in a smart and measured way. There is definite momentum the direction of sharing original research data and I’m looking forward to having more discussions about it on this blog.

Posted in openData | Leave a comment

The Problem with Paper Notebooks

The laboratory notebook is one of the most important tools for data management in the laboratory and in its paper format, it’s also one of the most problematic.

The paper laboratory notebook has historically been the place to record all of the information about an experiment: experimental data, experimental observations, the researcher’s thoughts, etc. Much of this is still true today, with the exception that most of our research data is digital and doesn’t fit nicely into a paper format. Instead, we print out tables and graphs and tape them into our notebooks as a bad approximation of a complete laboratory record.

Because of digital research, we’ve fundamentally divided our data from the document that records the context of that data. This is a problem because data without context are useless, as is context without the data. So we do our best to partner the two disparate systems of paper and electrons in order to have a usable laboratory record. This is frustrating, difficult to do well, and is having a major impact on the way we manage our data.

The best solution to the paper-digital divide is to fundamentally change the way we record information in the laboratory by using electronic lab notebooks. Having both the data and their context be digital and stored together will dramatically improve organization and searching. Additionally, e-lab notebook software is finally becoming viable and such systems are slowly being integrated into laboratories around the world. The benefits (and drawbacks) of e-lab notebooks require their own separate post, which I promise to write soon.

In the absence of an e-lab notebook, here are several suggestions for bridging the paper-digital divide:

  • Use one organization scheme. If your notebook is organized chronologically then your digital files should be organized in the same way. This will make it easier to find things.
  • Organize your digital files with respect to the notebook that they belong with. This may involve keeping a separate folder for each notebook you use.
  • Record the computer on which the digital files are stored (but remember that files may move).
  • Utilize indexes. You should definitely have one for your paper notebook and it would be useful to have an electronic version.
  • Keep your data with your lab notebook by writing all of the relevant digital files to disk and tucking that disk inside the cover of the paper notebook.
  • Digitally back up your paper notebooks by scanning them and storing the notes in the same folder as the data.

There are several other issues with paper lab notebooks (legibility, fragility, difficulty of searching, effort to back up), but the paper-digital divide is one of the biggest obstacles to to good data management in the laboratory. This problem is solvable by transitioning entirely to digital, but we need to be sure to do this in a smart way that ensures access to our data for years to come. In the meantime, small changes can make a big difference.

Posted in documentation, labNotebooks | Leave a comment

Who Owns My Data?

A recent post on the Retraction Watch blog–concerning a grad student retracting a sole-author paper when her advisor claimed ownership of her data–highlights the complicated nature of data ownership. This area is so complex that the only answer to the question of “who owns my data?” that I can provide is: it depends.

There are a lot of parties with an interest in your research data, not limited to: your funder, your institution, your boss/advisor, your collaborators, and you. Each of these entities may have a policy (explicit or not) on who has ownership of and who gets access to the data. My alma mater, for example, has a policy that establishes the university as the owner of the data but the PI as the steward of the data, meaning that PIs get to make most all of the decisions about data generated in their labs. Funders, on the other hand, may not exert ownership over your data but may instead require you to share it.

My best advice in this area is to assume that you don’t have ownership of your research data–especially if you are a grad student–until you look into your local policies. Even if you do all of the research, the entities who provide the equipment, laboratory space, and money may still lay claim to the data.

When it comes to data ownership, it’s much better to be conservative than to unintentionally burn bridges, end up with a retraction, or be just another bit of news  in the scientific blogosphere.

Posted in ownership | 1 Comment