Showing posts with label novelty. Show all posts
Showing posts with label novelty. Show all posts

Thursday, 24 September 2026

Trust in data: reproducibility packages do not guarantee validity

Note: This post builds on a series of PubPeer comments made here.

Last year the World Bank made news when it announced that all its outputs would include a reproducibility package, which contains all the data and code needed to reproduce the figures and tables in the report. In this regard it is following the lead of major journals in economics. Compared with some other fields, such as clinical trials (Avenell & Bishop, 2026), the widespread availability of data and code in economics publications is impressive.

Nevertheless, the confidence given by reproducibility packages needs to be tempered by one important consideration: a reproducibility package is only as good as the data that it is based on and the method of analysis that is selected. If the underlying dataset is flawed or the analysis isn't suited to the research question, then the conclusions may be unjustified, even if the whole analysis reproduces perfectly.

An example: the "edge science" study by Packalen and Bhattacharya(2020)

Before he became Director of the US National Institutes of Health (NIH), Jay Bhattacharya was professor of medicine, economics, and health research policy at Stanford University. In 2020, with co-author Mikko Packalen, he published a study in PNAS entitled "NIH funding and the pursuit of edge science". I first heard about this article when Dr Bhattacharya mentioned it in a workshop organised by the National Academies of Science, Engineering and Medicine on Enhancing Scientific Integrity: Progress and Opportunities in the Social and Behavioral Sciences. Specifically he said*:

"So a few years back I did a paper with Mikko Packalen where we estimated how old were the ideas in NIH published research. You go back to the way we measure the age of ideas - basically we just took all the words that were published in biomedicine in 1940 - word and word phrases - often [for] the 1940 word ... the idea [was] produced in 1941. You keep doing that you get a history of biomedicine. You go back to every paper you can figure out how old are the ideas are.  Turns out that papers by NIH funded researchers in the 1980s were relying on ideas ... zero, one or two years old. Whereas ... papers in ... 2017, they were relying on ideas that were like 7 or 8 years old. ARPA-H exists because there was a sense that the NIH wasn't funding research at the bleeding edge. It wasn't possible to do the kind of very innovative things through the NIH process."

I suspect many readers might query whether research based on new ideas is better than research based on older ideas, but if you are willing to accept that premise, then this seemed like a pretty neat study. As a proxy for “ideas", you get a bunch of words and phrases collated from scientific literature over an extended period. A pre-existing database, the Unified Medical Language System (UMLS) was used for this. For each phrase in the database you work out its cohort, i.e., the year when it first occurred in scientific literature.

Then you can take a set of articles and check which words occur in their titles and abstracts. That way you can check the novelty of the ideas in articles published at different times or by different funders. In this case the corpus consisted of all research articles by US researchers retrieved from the Medline database. For the time period 2010-2016 there were over 6 million publications, just under 1 million of which met the authors' criteria (research articles, US first author, at least 100 words in Title+Abstract). Packalen and Bhattacharya's (henceforth P&B) main research question was whether novelty of ideas in US-based biomedical research depended on whether or not the research was funded by NIH.

I am planning to deposit code and results for a detailed re-analysis of this study. For now,  I am just summarising the principal conclusions, focusing on the main result presented by the authors in Figure 1.

Problems with the dataset

The reproducibility package provided for the PNAS paper is deposited on the Harvard Dataverse. There is a set of "idea input" files for each year from 1990-2016, corresponding to each of 127 idea types. Each file has the PMID for each article and includes the idea type (category set), plus the most recent word corresponding to the idea type, and its vintage (referred to as cohort) - which is the first year when that word featured in the UMLS database. So the first few lines of data for category set 127 (Conceptual entity) in year 2010 look like this:

Table 1: First few rows of "idea input" file for category set 127
PMID Cohort Word Category set
17945460 1963 bulls eye 127
18161020 1953 customer 127
18264753 1951 adversity 127
18367639 1981 stakeholder 127
18372432 1953 customer 127
18455835 1961 diffusivity 127

When all category sets are combined, there's often more than one row for a single PMID, because a new row is created for each categoryset_number that is relevant for that article. For instance, PMID 18161020 includes the word "spouse" (vintage 1950), which is categorised under "Family Group" (category set 82) "Idea or Concept" (category set 89) and "Intellectual Product" (category set 114).

Another set of files (category status) provides information on PMID, journalcategoryid, year, and NIH_status for articles in a specific time period. For the analysis in Figure 1 information from both files was merged to get a file representing each article's contribution(s). 

The authors stress that their unit of analysis is the contribution rather than the article, so there are typically many contributions per article. To give an example of how this works, PMID 18161020 is published in the journal Aids and Behavior, which is categorised under two journal categories: Behavioral Sciences (Journal category 31) and Acquired Immunodeficiency Syndrome (Journal category 86). This article includes six words/phrases (“close friend”, “customer”, “multiple sexual partners”, “risk behaviors”, “selfefficacy”, “sexual partner”) that each belong to one category set (“Group”, “Conceptual Entity”, “Finding”, “Individual Behavior”, “Mental Process”, and “Population Group”, respectively), one phrase (“alcohol use”) that belongs to two category sets (“Clinical Attribute” and “Mental or Behavioral Dysfunction”), and one word (“spouse”) that belongs to the 3 category sets noted above. It thus is recorded as containing 22 contributions, 6 x 1 x 2 = 12, plus 1 x 2 x 2 = 4, plus 1 x 3 x 2 = 6.

This multiple counting of specific ideas (i.e. words/phrases) might seem reasonable if allocation of ideas to categories was straightforward, but I didn't find it so. Some idea categories such as "Finding", "Idea or Concept" or "Conceptual Entity" are extremely vague and I didn't understand the logic of why some phrases were assigned to them.

Furthermore, the coding of ideas for each article also seemed unintuitive. The title of PMID 18161020 is Multiple Sexual Partnerships in a Sample of African-American Crack Smokers, but I could not find the terms African-American, crack, cocaine or condom anywhere in the database. There was little overlap between the ideas coded by P&B and the MeSH codes allocated by PubMed.

This led me to check the database for terms from my own subject area.  One of my research topics, "handedness'", was not included in the set of ideas coded by P&B, so I was curious as to how items on this topic were coded. I checked PubMed for "handedness" in date range 2010-2016, and found over 15,000 articles. The first four research articles by US authors are shown in Table 2. All had "handedness" in the title. But PMID20943944 was categorised as "handedness right" in Idea category 124 (Organism attribute), PMID25204683 as "hand preference" in Idea category 36 (Finding), PMID20639549 as Edinburgh Handedness Inventory in Idea category 114 (Intellectual product), and PMID 23398763 had no code related to handedness. There was a 30 year difference between the idea cohorts for "hand preference" (1955) and "handedness right" (1985), but I could see no justification for these different codes.

Table 2: Coding for ideas related to "handedness" in 4 sample articles
Article Idea (Cohort) Category set
20639549: Handedness, heritability, neurocognition and brain asymmetry in schizophrenia edinburgh handedness inventory (1984) 114, Intellectual product
20943944: Handedness, dexterity, and motor cortical representations handedness right (1985) 124, Organism attribute
23398763: Handedness and corpus callosal morphology in Williams syndrome/ - -
25204683: Evaluating handedness measures in spider monkeys hand preference (1955) 36, Finding

It could be that these examples of counterintuitive codes are atypical, though I found I did not have to search to find these cases - a similar picture was found with other topics (e.g. "aphasia", "autism") - odd examples of idea coding popped up as soon as I searched PubMed for these topics. It's possible there is a coherent explanation, but it is difficult to evaluate this aspect of the study because insufficient information is given as to how the dataset was created. We know it was derived from public databases from MEDLINE and UMLS, but it remains unclear precisely how these two sources of information were combined.

Problems with the method of analysis

I also found that results depended crucially on how the dependent variable, novelty of ideas, was defined. It turns out there are various ways you can quantify this, and the choice of analysis has a big impact on the results.

P&B's main result, which I replicated, was presented thus: 

Figure 1 from Packalen and Bhattacharya (2020) 



I encourage readers to look carefully at this Figure to check if you can understand it. It looks simple but has some puzzling features.

The clear thing that springs out from the plot is that the seven points on the right hand end of the y-axis fall below the dotted line, which corresponds to the mean rate of NIH funding overall for this set of US-based research studies.

This result is the basis for Bhattacharya's statement "papers in 2017 were relying on ideas that were like 7 or 8 years old", echoed in the Abstract of P&B (2020) with the statement

"Papers that build on very recent ideas are NIH funded less often than are papers that build on ideas that have had a chance to mature for at least 7 y. "

There are two odd aspects of this analysis. The first is the x-axis, showing "cohort of newest idea input". Figure 1 is based on all 992,633 US-based biomedical research papers published during 2010 to 2016, but the date of publication is not considered. This makes little sense, because a paper cannot include an idea that is more recent than its publication date. Furthermore, an idea of cohort 2010 would be new if included in a paper published in 2011, but relatively old in a paper published in 2016. Such subtleties are missed in Figure 1.

Second, papers are not represented in Figure 1. As noted above, what you are seeing are "contributions". As the authors explain "We count a paper that is linked to K idea types and J research areas as K*J contributions". 

Because Figure 1 looks at contributions rather than papers, it misses the key point about many papers, which is that their contribution represents a combination of ideas. Suppose an NIH-funded article reported a new treatment (cohort 2016) applied to five well-known diseases (cohort in 1950s) - this would contribute five datapoints for old ideas to Figure 1 and only one for a novel idea.

It struck me that a better way of indexing novelty of ideas in NIH-funded research would be to treat the paper as the unit of analysis, and identify the most recent idea included in a paper, measured as the smallest gap between the cohort of the idea and the year of publication. If you do that, and plot the data as in Figure 1, you get Figure 2. 

Figure 2: Plot as for Figure 1, but with paper as the unit of analysis, where x-axis shows gap between idea 'cohort' and publication year of the paper featuring that idea


As you can see, the evidence now goes directly against the conclusion "Papers that build on very recent ideas are NIH funded less often than are papers that build on ideas that have had a chance to mature for at least 7 y." The points for a gap of 1 to 7 years are above the dotted line - the point for zero years is the only one slightly below the line. So, it now looks as if ideas from the past 7 years are more likely to be funded by NIH than by other funders.

It's always concerning when a modification of the analysis reverses the outcome, and I contacted the authors to ask what they made of this.

Dr Packalen gave a detailed reply, referring to the famous "garden of forking paths" and the numerous decisions that have to be made at each step of analysis that potentially would allow for thousands of analyses. He mentioned a number of factors that might affect which analysis was selected, including pressure from the academic community to obtain a specific result, and ease of interpretation of the results.

Neither of these satisfied me, because I had not found the published analysis to be easy to interpret - quite the opposite, and I was concerned that if a change in analysis could flip the finding from NIH looking weak on edge science to it looking strong on edge science, then this would merit some discussion.

Regarding the reliance on "contributions" rather than articles, Dr Packalen replied that this was a consequence of rejecting "the Great Man Theory of Science".  But an analysis using the paper as the unit of analysis does not entail we believe in the Great Man Theory of Science - it just means that an article contributes to knowledge with a constellation of ideas that come together, and that to separate them is like evaluating the deliciousness of a cake from a list of its ingredients.

In a previous article, Packalen and Bhattacharya (2017) had addressed similar issues from their large database to ask which journals published most novel ideas. I pointed out that in that article, they adopted the analytic approach I used for Figure 2, i.e. focusing on the gap between publication date and cohort, and taking the shortest gap for each article. I asked why they had abandoned this approach - which seemed more logical to me - in the 2020 paper.

Dr Packalen replied saying that both the decision to base the analysis on contributions rather than papers, and to focus on absolute date of ideas, rather than novelty relative to publication year, were "forks", implying that some degree of arbitrariness is to be expected in such an analysis. Indeed,  the analyses included in the reproducibility package are numbered 66, 88, 89, 90, 91 and 93, suggesting there were many more unreported analyses. Dr Packalen pointed out that some alternative analyses had been provided in supplementary material. While this is the case, these analyses all used idea cohort on the x-axis, and as far as I could see, none looked at the gap between the idea cohort and publication date (as in Figure 2).

The impression was that he was quite relaxed about accepting that there were many possible analyses and they might give different answers, implying that perhaps it was naive to imagine there could be one true analysis. I suspect this viewpoint may be normative in his field.  But, as Opoku-Agyemang pointed out in this blogpost 

"While it is encouraging that a growing number of journals are requiring authors to provide all data and code, the conference audience agreed that the code submitted to a journal typically represents a 'selection bias' of sorts, relative to the universe of queries actually implemented on the data. Thus, just replicating what authors provide is a frightfully low bar for transparency because we never know what non-submitted queries were run on a dataset and how these queries may have changed the outcome of the research."

Given my background in psychology, I'm not at all relaxed about wandering around the garden of forking paths and selecting your analysis after inspecting the results. We know that approach leads to non-reproducible studies and much has been written about how to overcome the resulting bias. In experimental disciplines, one solution is to adopt the Registered Reports approach, where you preregister your methods and analysis plan, and have your paper reviewed prior to results being collected. 

If you are analysing an existing dataset, that's not feasible (unless you can prove you have not peeked at the data prior to developing an analysis plan). In such a situation some might recommend conducting a range of analyses and reporting them all - in the extreme case doing a multiverse analysis. However, as Lakens et al (2026) have argued, different approaches to analysis entail different auxiliary assumptions about the link between theory and empirical observations; when a multiverse analysis reveals different consequences of analytic decisions, this may help clarify what auxiliary assumptions are involved, but at the end of the day, some assumptions are theoretically justified and others are not. In a similar vein, Bishop & Hulme (2024) looked at a real life example of alternative analyses of an educational trial, and rejected the idea that all alternative analyses are equally valid. We suggested that a good way to clarify underlying theoretical assumptions is by using simulation to generate data from a model, and then testing how well different analytic approaches captured the underlying model. It would be fascinating to try this approach with the edge science data.

In my view it is not defensible to draw conclusions about the novelty of ideas in published articles by using a dependent variable that does not use the article as the unit of analysis, and that neglects to take into account the gap between publication date and idea vintage. This seems epistemically incoherent.  The problem is made worse by the lack of clarity of how ideas were assigned to articles.

Having said all that, there's no doubt that having a Reproducibility Package is a huge advance on the times when no data was made public.  It has made it possible for me to evaluate the dataset and the analytic approach in a way that would have been otherwise impossible, and I am grateful to the authors for making the materials available and engaging with my queries**. 

However, I fear that if we combine a relaxed approach to researcher choice of forking paths with a stringent requirement for open data and code, we run the risk of giving unwarranted credibility to analytic decisions that may be non-optimal at best and positively misleading at worst. This is particularly concerning when results are used to inform major policy decisions. 

Notes 

*The full unedited text is appended to an earlier blogpost 

**I have paraphrased the arguments used in Prof Packalen's email; I have requested his permission to reproduce the email in full, and will append it here if this is given. 

As usual, I welcome on-topic comments, but seldom post them if anonymous. Comments are moderated to prevent spam and so there is typically a lag of a few days before they appear. 

Monday, 27 October 2025

Problems with ELife's new article type: Replication studies

 

I was interested to receive an email from eLife last week, telling me that "As part of our commitment to open science, scientific rigour and transparency we now accept submissions of Replication Studies". Sounds good, I thought, but on reading further I became increasingly dismayed. I think the way this is set up dooms it to failure.

Back in 2023, when eLife created a storm by altering its publishing model, I argued that they would not achieve their aims unless they changed the basis for selecting articles for peer review.  Although eLife has abandoned the traditional binary outcome where articles are accepted or rejected, there is still a decision delegated to editors, which is whether the article is selected for peer review. My suggestion was that they should adopt results-blind selection. This avoids the bias that favours articles with positive results. Publication bias leaves well-conducted studies with null results to languish unpublished. This matters because a cumulative science should include negative as well as positive findings. Nissen et al (2016) showed how publication bias leads to what they termed the "canonization of false facts". It is a universal phenomenon and I doubt that eLife editors are immune to it.

One method of avoiding publication bias is the Registered Reports format, whereby the introduction, methods and analysis plan are peer reviewed before data are collected, with the article accepted in principle by a journal, provided the researchers did the study as planned, or gave adequate reasons for deviating from protocol. In my experience, biomedical researchers are highly resistant to Registered Reports; they tell me the approach is incompatible with how they work, which involves development of ideas and methods in the course of doing a study 🙄. However, that argument does not hold for replication studies, where the idea is to reproduce the methods of an existing published study. Indeed, eLife took a pioneering stand in hosting the Reproducibility Project on Cancer Biology (Errington et al., 2021a). So it is disappointing that their current proposal is for a much more timid approach.

First of all, it sounds as if the plan is for individuals to complete a replication study before submitting  it to eLife. This means that we'll be up against publication bias all over again - the likelihood of an editor selecting a study for peer review will depend not on the strength and suitability of methods but on whether the replication was "successful".

Second the eLife instructions state: "The authors should work closely with the authors of the original study and summarize their interactions with the original authors as part of a cover letter". It seems entirely appropriate to liaise with original authors to ensure that materials and methods are suitable for replication. However, we know from the Cancer Biology Replication project that many authors were unresponsive or even obstructive when asked to give advice on a replication study (Errington et al, 2021b). Many attempts at replication failed because it was impossible to work out exactly what had been done in the original study or because authors would not share materials. In effect, then, the current eLife approach to replication studies gives original authors the ability to veto any replication attempt.

I can't think who on earth would take the risk of doing a replication study under these conditions. Many scientists already regard replication as an inferior type of research activity, for which replicators get little credit - and potential abuse. Furthermore, it is difficult to get funds for replication studies, because they are deemed insufficiently novel. Yet, previous large-scale replication studies have found that direct replications often fail to find original effects, or replicate the effect but with a smaller effect size. So you make yourself unpopular by attempting a replication, and then when you come to try to submit it to eLife you are told it won't be peer reviewed because you obtained a null effect.

I think that we won't get replication studies in biosciences unless they are explicitly incentivised - and judged on their methodological quality rather than their results. Meanwhile, because of the field's obsession with novelty, research progress stalls as people keep trying to build on results without knowing whether they provide a solid foundation.

My prediction: give it a year, and see how many Replication Studies have been accepted for peer review in eLife. I'll be surprised if it's more than zero.

References

Errington, T. M., Mathur, M., Soderberg, C. K., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021a). Investigating the replicability of preclinical cancer biology. eLife, 10, e71601. https://doi.org/10.7554/eLife.71601

Errington, T. M., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021). Reproducibility in Cancer Biology: Challenges for assessing replicability in preclinical cancer biology | eLife. eLife. https://doi.org/10.7554/eLife.67995

Nissen, S. B., Magidson, T., Gross, K., & Bergstrom, C. T. (2016). Publication bias and the canonization of false facts. eLife, 5, e21451. https://doi.org/10.7554/eLife.21451