Showing posts with label open data. Show all posts
Showing posts with label open data. Show all posts

Sunday, 26 May 2024

Are commitments to open data policies worth the paper they are written on?

 

As Betteridge's law of headlines states: "Any headline that ends in a question mark can be answered by the word no."  So you know where I am going with this.  

 

I'm a longstanding fan of open data - in fact, I first blogged about this back in 2015. So I've been gratified to see the needle shift on this, in the sense that over the past decade, in a rush to present themselves as good guys, various institutions and publishers have published policies supporting open data. The problem is that when you actually ask them to implement those policies, they back down.   

 

I discussed arguments for and against data-sharing in a Commentary article in 2016. I divided the issues according to whether they focused on the impact of data-sharing on researchers or on research participants. Table 1 from that article, entitled "Conflict between interests of researchers and advancement of science" is reproduced here:

 

Argument

Counter-argument

1. Lack of time to curate data.

Unless adequately curated, data will over time become unusable, including by the original researcher.

2. Personal investment—reluctance to give data to freeloaders.

Reuse of data increases its value and the researcher benefits from additional citations. There is also an ethical case for maximizing use of data obtained via public funding.

3. Concerns about being scooped before the analysis is complete.

This is a common concern though there are few attested cases. A time-limited period of privileged use by the study team can be specified to avoid scooping.

4. Fear of errors being found in the data.

Culture change is needed to recognize errors are inevitable in any large dataset and should not be a reason for reputational damage. Data-sharing allows errors to be found and corrected.

 

I then went on to discuss two other concerns which focused on implications of data-sharing for human participants, viz:

5.  Ethical concerns about confidentiality of personal data, especially in the context of clinical research

6.  Possibility that others with a different agenda may misuse the data, e.g. perform selective analyses that misrepresent the findings.

 

These last two issues raise complex concerns and there's plenty to discuss on how address them, but I'll put that to one side for now, as the case I want to comment on concerns a simple dataset where there is limited scope for secondary analyses and where no human participants are involved.

 

My interest was piqued by comments on PubPeer about a paper entitled "Magnetic field screening in hydrogen-rich high-temperature superconductors ".  The thread on PubPeer starts with this extraordinary comment by J. E. Hirsch:

 

I requested the underlying data for Figs. 3a, 3e, 3b, 3f of this paper on Jan 11, 2023. This is because the published data for Figs. 3a and 3e, as well as for Figs. 3b and 3f, are nominally the same but incompatible with each other, and I would like to understand why that is. I asked the authors to explain, but they did not provide an explanation. Neither did they supply the data. The journal told me that it had received the data from the authors but will not share them with me because they are "confidential". I requested that the journal posts an Editor Note informing readers that data are unavailable to readers. The journal responded that because data were share with editors they "cannot write an editorial note on the published article stating the data is unavailable as this would be factually incorrect".

 

Pseudonymous commenter Orchestes quercus drew attention to the Data Availability statement in the article: "The data that support the findings of this study are available from the corresponding authors upon reasonable request".

 

J. E. Hirsch then added a further comment: 

 

The underlying data are still not available, the editor says the author deems the request "unreasonable" but it cannot divulge the reasoning behind it, nor can the journal publish an editor note that there are restrictions on data availability because the data were provided to the journal.  Springer Nature's Research Integrity Director wrote to me in September 2023 that "we recognize the right of the authors to not share the data with you, in line with the authors’ chosen data availability statement", and that "As Springer Nature considers the correspondence with the authors confidential, we cannot share with you any further details.

 

Now, I know nothing whatsoever about superconductors or J. E. Hirsch, but I think the editors, publisher and the authors are making themselves look very silly, and indeed suspicious, by refusing to share the data.  They can't plead patient confidentiality or ethical restrictions - it seems they are just refusing to comply because they don't want to.  

 

To up the ante, Orchestes quercus extracted data from the figures and did further analyses, which confirmed that J. E. Hirsch had a point - the data did not appear to be internally consistent.

 

Meanwhile, I had joined the PubPeer thread, pointing out

 

The authors and editor appear to be in breach of the policy of Nature Portfolio journals, stated here:
https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards, viz:

An inherent principle of publication is that others should be able to replicate and build upon the authors' published claims. A condition of publication in a Nature Portfolio journal is that authors are required to make materials, data, code, and associated protocols promptly available to readers without undue qualifications. Any restrictions on the availability of materials or information must be disclosed to the editors at the time of submission. Any restrictions must also be disclosed in the submitted manuscript.

After publication, readers who encounter refusal by the authors to comply with these policies should contact the chief editor of the journal. In cases where editors are unable to resolve a complaint, the journal may refer the matter to the authors' funding institution and/or publish a formal statement of correction, attached online to the publication, stating that readers have been unable to obtain necessary materials to replicate the findings.

I also noted that two of the authors are based at a Max Planck Institute. The Max Planck Gesellschaft is a signatory to the BerlinDeclaration on Open Access to Knowledge in the Sciences and Humanities.  On the website it states:

the Max Planck Society (MPG) is committed to the goal of providing free and open access to all publications and data from scholarly research (my emphasis).

 

Well, the redoubtable J. E. Hirsch had already thought of that, and in a subsequent PubPeer comment made public various exchanges he had had with luminaries from the Max Planck Institutes.

 

All I can say to the Max Planck Gesellschaft is that this is not a good look. Hirsch has noted an inconsistency in the published figures.  This has been confirmed by another reader and needs to be explained. The longer people dig in defensively, attacking the person making the request rather than just showing the raw data, the more it looks as if something fishy is going on here.

 

Why am I so hung up on data-sharing? The reason is simple. The more I share my own data, or use data shared by others, the more I appreciate the value of doing so. Errors are ubiquitous, even when researchers are careful, but we'll never know about them if data are locked away.

 

Furthermore, it is a sad reality that fraudulent papers are on the rise, and open data is one way of defending against them. It's not a perfect defence: people can invent raw data as well as summary data, but realistic data are not so easy to fake, and requiring open data would slow down the fraudsters and make them easier to catch.

 

Having said that, asking for data is not tantamount to accusing researchers of fraud: it should be accepted as normal scientific practice to make data available in order that others can check the reproducibility of findings. If someone treats such a request as an accusation, or deems it "unreasonable", then I'm afraid it just makes me suspicious.  

 

And if organisations like Springer Nature and Max Planck Gesellschaft won't back up their policies with action, then I think they should delete them from their websites. They are presenting themselves as champions of open, reproducible science, while acting as defenders of non-transparent, secret practices. As we say in the UK, fine words butter no parsnips.   

 

 P.S. 27th May:  A comprehensive account of the superconductivity affair has just appeared on the website For Better Science.  This suggests things are even worse than I thought.   

 

In addition, you can see Jorge Hirsch explain his arduous journey in attempting to access the data here.    

 

NOTE ON COMMENTS: Many thanks to those who have commented. Comments are moderated to prevent spam, so there is a delay before they appear, but I will accept on-topic comments in due course.




Wednesday, 9 May 2018

My response to the EPA's 'Strengthening Transparency in Regulatory Science'

Incredible things have happened at the US Environmental Protection Agency since Donald Trump was elected. The agency is responsible for creating standards and laws that promote the health of individuals and the environment. During previous administrations it has overseen laws concerned with controlling pollution and regulating carbon emissions. Now, under Administrator Scott Pruitt, the voice of industry and climate scepticism is in the ascendant. 

A new rule that purports to 'Strengthen Transparency in Regulatory Science' has now been proposed - ironically, at a time when the EPA is being accused of a culture of secrecy regarding its own inner workings. Anyone can comment on the rule here: I have done so, but my comment appears to be in moderation, so I am posting it here.


Dear Mr Pruitt

re: Regulatory Science- Docket ID No. EPA-HQ-OA-2018-0259

The proposed rule, ‘Strengthening transparency in regulatory science’ brings together two strands of contemporary scientific activity. On the one hand, there is a trend to make policy more evidence-based and transparent. On the other hand, there has, over the past decade, been growing awareness of problems with how science is being done, leading to research that is not always reproducible (the same results achieved by re-analysis of the data) or replicable (similar results when an experiment is repeated). The proposed rule by the Environmental Protection Agency (EPA) brings these two strands together by proposing that policy should only be based on research that has openly available public data. While this may on the surface sound like a progressive way of integrating these two strands, it rests on an over-simplified view of how science works and has considerable potential for doing harm.

I am writing in a personal capacity, as someone at the forefront of moves to improve reproducibility and replication of science in the UK. I chaired a symposium at the Academy of Medical Sciences on this topic in 2015; this was jointly organised with UK major funders: Wellcome Trust, Medical Research Council and Biotechnology and Biological Science Research Council (https://acmedsci.ac.uk/policy/policy-projects/reproducibility-and-reliability-of-biomedical-research). I am involved in training early career researchers in methods to improve reproducibility, and I am a co-author of Munafò, M. R et al  (2017). A manifesto for reproducible science. Nature Human Behavior, 1(1: 0021). doi:10.1038/s41562-016-0021. I would welcome any move by the US Government that would strengthen research by encouraging adoption of methods to improve science, including making analysis scripts and data open when this is not in conflict with legal/ethical issues. Unfortunately, this proposal will not do that. Instead, it will weaken science by drawing a line in the sand that effectively disregards scientific discoveries prior to the first part of the 21st century when the importance of open data started to be increasingly recognised.

The proposal ignores a key point about how scientific research works: the time scale. Most studies that would be relevant to EPA take years to do, and even longer to filter through to affect policy. Consequences for people and the planet that are relevant to environmental protection are often not immediately obvious: if they were, we would not need research. Recognition, for instance, of the dangers of asbestos, took years because the impacts on health were not immediate. Work demonstrating the connection occurred many years ago and I doubt that the data are anywhere openly available, yet the EPA’s proposed rule would imply that it could be disregarded. Similarly, I doubt there is open data demonstrating the impact of lead in paint or exhaust fumes, or of pesticides such as DDT: does this mean that manufacturers would be free to reintroduce these?

A second point is that scientific advances never depend on a single study: having open scripts and data is one way of improving our ability to check findings, but it is a relatively recent development, and it is certainly not the only way to validate science. The growth of knowledge has always depended on converging evidence from different sources, replications by different scientists and theoretical understanding of mechanisms. Scientific facts become established when the evidence is overwhelming. The EPA proposal would throw out the mass of accumulated scientific evidence from the past, when open practices were not customary – and indeed often not practical before computers for big data were available.

Contemporary scientific research is far from perfect, but the solution is not to ignore it, but to take steps to improve it and to educate policy-makers in how to identify strong science; government needs advisors who have scientific expertise and no conflict of interest, who can integrate existing evidence with policy implications. The ‘Strengthening Transparency’ proposal is short-sighted and dangerous and appears to have been developed by people with little understanding of science. It puts citizens at risk of significant damage – both to health and prosperity -  and it will make the US look scientifically illiterate to the rest of the world.


Yours sincerely


D. V. M. Bishop FMedSci, FBA, FRS


Sunday, 15 November 2015

Who's afraid of Open Data

Cartoon by John R. McKiernan, downloaded from: http://whyopenresearch.org/gallery.html

I was at a small conference last year, catching up on gossip over drinks, and somehow the topic moved on to journals, and the pros and cons of publishing in different outlets. I was doing my best to advocate for open access, and to challenge the obsession with journal impact factors. I was getting the usual stuff about how early-career scientists couldn't hope to have a career unless they had papers in Nature and Science, but then the conversation took an interesting turn.

"Anyhow," said eminent Professor X. "One of my postdocs had a really bad experience with a PLOS journal."

Everyone was agog. Nothing better at conference drinks than a new twist on the story of evil reviewer 3.  We waited for him to continue. But the problem was not with the reviewers.

"Yup. She published this paper in PLOS Biology, and of course she signed all their forms. She then gave a talk about the study, and there was this man in the audience, someone from a poky little university that nobody had ever heard of, who started challenging her conclusions. She debated with him, but then, when she gets back she has an email from him asking for her data."

We wait with bated breath for the next revelation.

"Well, she refused of course, but then this despicable person wrote to the journal, and they told her that she had to give it to him! It was in the papers she had signed."

Murmurs of sympathy from those gathered round. Except, of course, me. I just waited for the denouement. What had happened next, I asked.

"She had to give him the data. It was really terrible. I mean, she's just a young researcher starting out."

I was still waiting for the denouement. Except that there was no more. That was it! Being made to give your data to someone was a terrible thing. So, being me, I asked, why was that a problem? Several people looked at me as if I was crazy.

"Well, how would you like it if you had spent years of your life gathering data, data which you might want to analyse further, and some person you have never heard comes out of nowhere demanding to have it?"

"Well, they won't stop you analysing it," I said.

"But they may scoop you and find something interesting in it before you have a chance to publish it!"

I was reminded of all of this at a small meeting that we had in Oxford last week, following up on the publication of a report of a symposium I'd chaired on Reproducibility and Reliability of Biomedical Research. Thanks to funding by the St John's College Research Centre, a small group of us were able to get together to consider ways in which we could take forward some of the ideas in the report for enhancing reproducibility. We covered a number of topics, but the one I want to focus on here is data-sharing.

A move toward making data and analyses open is being promoted in a top-down fashion by several journals, and universities and publishers have been developing platforms to make this possible. But many scientists are resisting this process, and putting forward all kinds of argument against it. I think we have to take such concerns seriously: it is all too easy to mandate new actions for scientists to follow that have unintended consequences and just lead to time-wasting, bureaucracy or perverse incentives. But in this case I don't think the objections withstand scrutiny. Here are the main ones we identified at our meeting:

1.  Lack of time to curate data;  Data are only useful if they are understandable, and documenting a dataset adequately is a non-trivial task;

2.  Personal investment - sense of not wanting to give away data that had taken time and trouble to collect to other researchers who are perceived as freeloaders;

3. Concerns about being scooped before the analysis is complete;

4.  Fear of errors being found in the data;

5.  Ethical concerns about confidentiality of personal data, especially in the context of clinical research;

6.  Possibility that others with a different agenda may misuse the data, e.g. perform selective analysis that misrepresented the findings;

These have partial overlap with points raised by Giorgio Ascoli (2015) when describing NeuroMorpho.Org, an online data-sharing repository for digital reconstructions of neuronal morphology. Despite the great success of the repository, it is still the case that many people fail to respond to requests to share their data, and points 1 and 2 seemed the most common reasons.

As Ascoli noted, however, there are huge benefits to data-sharing, which outweigh the time costs. Shared data can be used for studies that go beyond the scope of the original work, with particular benefits arising when there is pooling of datasets. Some illustrative examples from the field of brain imaging were provided by Thomas Nichols at our meeting (slides here), where a range of initiatives is being developed to facilitate open data. Data-sharing is also beneficial for reproducibility: researchers will check data more carefully when it is to be shared, and even if nobody consults the data, the fact it is available gives confidence in the findings. Shared data can also be invaluable for hands-on training. A nice example comes from Nicole Janz, who teaches a replication workshop in social sciences in Cambridge, where students pick a recently published article in their field and try to obtain the data so they can replicate the analysis and results.

These are mostly benefits to the scientific community, but what about the 'freeloader' argument? Why should others benefit when you have done all the hard work? In fact, when we consider that scientists are usually receiving public money to make scientific discoveries, this line of argument does not appear morally defensible. But in any case, it is not true that the scientists who do the sharing have no benefits. For a start, they will see an increase in citations, as others use their data. And another point, often overlooked, is that uncurated data often become unusable by the original researcher, let alone other scientists, if it is not documented properly and stored on a safe digital site. Like many others, I've had the irritating experience of going back to some old data only to find I can't remember what some of the variable names refer to, or whether I should be focusing on the version called final, finalfinal, or ultimate. I've also had the experience of data being stored on a kind of floppy disk, or coded by a software package that had a brief flowering of life for around 5 years before disappearing completely.

Concerns about being scooped are frequently cited, but are seldom justified. Indeed, if we move to a situation where a dataset is a publication with its own identifier, then the original researcher will get credit every time someone else uses the dataset. And in general, having more than one person doing an analysis is an important safeguard, ensuring that results are truly replicable and not just a consequence of a particular analytic decision (see this article for an illustration of how re-analysis can change conclusions).

The 'fear of errors' argument is, of understandable but not defensible. The way to respond is to say of course there will be errors – there always are. We have to change our culture so that we do not regard it as a source of shame to publish data in which there are errors, but rather as an inevitability that is best dealt with by making the data public so the errors can be tracked down.

Ethical concerns about confidentiality of personal data are a different matter. In some cases, participants in a study have been given explicit reassurances that their data will not be shared: this was standard practice for many years before it was recognised that such blanket restrictions were unhelpful and typically went way beyond what most participants wanted – which was that their identifiable data would not be shared.  With training in sophisticated anonymization procedures, it is usually possible to create a dataset that can be shared safely without any risk to the privacy of personal information; researchers should be anticipating such usage and ensuring that participants are given the option to sign up to it.

Fears about misuse of data can be well-justified when researchers are working on controversial areas where they are subject to concerted attacks by groups with vested interests or ideological objections to their work. There are some instructive examples here and here. Nevertheless, my view is that such threats are best dealt with by making the data totally open. If this is done, any attempt to cherrypick or distort the results will be evident to any reputable scientist who scrutinises the data. This can take time and energy, but ultimately an unscientific attempt to discredit a scientist by alternative analysis will rebound on those who make it.  In that regard, science really is self-correcting. If the data are available, then different analyses may give different results, but a consensus of the competent should emerge in the long run, leaving the valid conclusions stronger than before.

I'd welcome comments from those who have started to use open data and to hear your experiences, good or bad.

P.S. As I was finalising this post, I came across some great tweets from the OpenCon meeting taking place right now in Brussels. Anyone seeking inspiration and guidance for moving to an open science model should follow the #opencon hashtag, which links to materials such as these: Slides from keynote by Erin McKiernan, and resources at http://whyopenresearch.org/about.html.

Monday, 26 May 2014

Data sharing: Exciting but scary

© www.CartoonStock.com

Yesterday I did something I've never done before in many  years of publishing. When I submitted a revised manuscript of a research report to a journal, I also posted the dataset on the web, together with the script I'd used to extract the summary results. It was exciting. It felt as if I was part of a scientific revolution that has been gathering pace over the past two or three years, which culminated in adoption of a data policy by PLOS journals last February. This specified that authors were required to make the data underlying their scientific findings available publicly immediately upon publication of the article. As it happens, my paper is not submitted to PLOS, and so I'm not obliged to do this, but I wanted to, having considered the pros and cons. My decision was also influenced by the Wellcome Trust, who fund my work and encourage data sharing.

The benefits are potentially huge. People usually think about the value to other researchers, who may be able to extract useful information from your data, and there's no doubt this is a factor.  Particularly with large datasets, it's often the case that researchers only use a subset of the data, and so valuable information is squandered and may be lost forever.  More than once I've had someone ask me for an old dataset, only to find it is inaccessible, because it was stored on a floppy disk or an ancient, non-networked computer and so is no longer readable.  Even if you think that you've extracted all you can from a dataset, it may still be worth preserving for potential inclusion in future meta-analyses.

Another value of open data is less often emphasised: when you share data you are forced to ensure it is accurate and properly documented. I enjoy data analysis, but I'm not naturally well-disciplined about keeping everything tidy and well-organised. I've been alarmed on occasion to return to a dataset and find I have no idea what some of the variables are, because I failed to document them properly.  If I know the world at large will see my dataset then I won't want to be embarrassed by it, and so I will take more care to keep it neat and tidy with everything clearly labelled. This can only be good.

But here's the scary thing. data sharing exposes researchers to the risk of being found out to be sloppy or inaccurate. To my horror, shortly before I posted my dataset on the internet yesterday I found I'd made a mistake in the calculation of one of my variables. It was a silly error, caused by basing a computation on the wrong column of data. Fortunately, it did not have a serious effect on my paper, though I did have to go through redoing all the tables and making some changes to the text.  But it seemed like pure chance that I picked up on this error – I could very easily have posted the dataset on the internet with the error still there. And it was an error that would have been detected by anyone eagle-eyed enough to look at the numbers carefully.  Needless to say, I'm nervous that there may well be other errors in there that I did not pick up. But at least it's not as bad as an apocryphal case of a distinguished research group whose dramatic (and published) results arose because someone forgot to designate 9 as a missing value code. When I heard about that I shuddered, as I could see how easily it could happen.

This is why Open Data is both important for science but difficult for scientists. In the past, I've found mistakes in my datasets, but this has been a private experience.  To date, as far as I am aware, no serious errors have got into my published papers – though I did have another close shave last year when I found a wrongly-reported set of means at the proofs stage, and there have been a couple of instances where minor errata have had to be published. But the one thing I've learned as I wiped the egg off my face is that error is inevitable and unavoidable, however careful you try to be. The best way to flush out these errors is to make the data public. This will inevitably lead to some embarrassment when mistakes are found, but at the end of the day, our goal must be to find out what is the case, rather than to save face.

I'm aware that not everyone agrees with me on this. There are concerns that open data sharing could lead to scientists getting scooped, will take up too much time, and could be used to impose ever more draconian regulation on beleaguered scientists: as DrugMonkey memorably put it:  "Data depository obsession gets us a little closer to home because the psychotics are the Open Access Eleventy waccaloons who, presumably, started out as nice, normal, reasonable scientists." But I think this misses the point. Drug Monkey seems to think this is all about imposing regulations to prevent fraud and other dubious practices.  I don't think this is so. The counter-arguments were well articulated in a blogpost by Tal Yarkoni. In brief, it's about moving to a point where it is accepted practice to make data publicly available, to improve scientific transparency, accuracy and collaboration.