Last year the World Bank made news when it announced that all its outputs would include a reproducibility package, which contains all the data and code needed to reproduce the figures and tables in the report. In this regard it is following the lead of major journals in economics. Compared with some other fields, such as clinical trials (Avenell & Bishop, 2026), the widespread availability of data and code in economics publications is impressive.
Nevertheless, the confidence given by reproducibility packages needs to be tempered by one important consideration: a reproducibility package is only as good as the data that it is based on and the method of analysis that is selected. If the underlying dataset is flawed or the analysis isn't suited to the research question, then the conclusions may be unjustified, even if the whole analysis reproduces perfectly.
An example: the "edge science" study by Packalen and Bhattacharya(2020)
Before he became Director of the US National Institutes of Health (NIH), Jay Bhattacharya was professor of medicine, economics, and health research policy at Stanford University. In 2020, with co-author Mikko Packalen, he published a study in PNAS entitled "NIH funding and the pursuit of edge science". I first heard about this article when Dr Bhattacharya mentioned it in a workshop organised by the National Academies of Science, Engineering and Medicine on Enhancing Scientific Integrity: Progress and Opportunities in the Social and Behavioral Sciences. Specifically he said*:
"So a few years back I did a paper with Mikko Packalen where we estimated how old were the ideas in NIH published research. You go back to the way we measure the age of ideas - basically we just took all the words that were published in biomedicine in 1940 - word and word phrases - often [for] the 1940 word ... the idea [was] produced in 1941. You keep doing that you get a history of biomedicine. You go back to every paper you can figure out how old are the ideas are. Turns out that papers by NIH funded researchers in the 1980s were relying on ideas ... zero, one or two years old. Whereas ... papers in ... 2017, they were relying on ideas that were like 7 or 8 years old. ARPA-H exists because there was a sense that the NIH wasn't funding research at the bleeding edge. It wasn't possible to do the kind of very innovative things through the NIH process."
I suspect many readers might query whether research based on new ideas is better than research based on older ideas, but if you are willing to accept that premise, then this seemed like a pretty neat study. As a proxy for “ideas", you get a bunch of words and phrases collated from scientific literature over an extended period. A pre-existing database, the Unified Medical Language System (UMLS) was used for this. For each phrase in the database you work out its cohort, i.e., the year when it first occurred in scientific literature.
Then you can take a set of articles and check which words occur in their titles and abstracts. That way you can check the novelty of the ideas in articles published at different times or by different funders. In this case the corpus consisted of all research articles by US researchers retrieved from the Medline database. For the time period 2010-2016 there were over 6 million publications, just under 1 million of which met the authors' criteria (research articles, US first author, at least 100 words in Title+Abstract). Packalen and Bhattacharya's (henceforth P&B) main research question was whether novelty of ideas in US-based biomedical research depended on whether or not the research was funded by NIH.
I am planning to deposit code and results for a detailed re-analysis of this study. For now, I am just summarising the principal conclusions, focusing on the main result presented by the authors in Figure 1.
Problems with the dataset
The reproducibility package provided for the PNAS paper is deposited on the Harvard Dataverse. There is a set of "idea input" files for each year from 1990-2016, corresponding to each of 127 idea types. Each file has the PMID for each article and includes the idea type (category set), plus the most recent word corresponding to the idea type, and its vintage (referred to as cohort) - which is the first year when that word featured in the UMLS database. So the first few lines of data for category set 127 (Conceptual entity) in year 2010 look like this:
| PMID | Cohort | Word | Category set |
|---|---|---|---|
| 17945460 | 1963 | bulls eye | 127 |
| 18161020 | 1953 | customer | 127 |
| 18264753 | 1951 | adversity | 127 |
| 18367639 | 1981 | stakeholder | 127 |
| 18372432 | 1953 | customer | 127 |
| 18455835 | 1961 | diffusivity | 127 |
When all category sets are combined, there's often more than one row for a single PMID, because a new row is created for each categoryset_number that is relevant for that article. For instance, PMID 18161020 includes the word "spouse" (vintage 1950), which is categorised under "Family Group" (category set 82) "Idea or Concept" (category set 89) and "Intellectual Product" (category set 114).
Another set of files (category status) provides information on PMID, journalcategoryid, year, and NIH_status for articles in a specific time period. For the analysis in Figure 1 information from both files was merged to get a file representing each article's contribution(s).
The authors stress that their unit of analysis is the contribution rather than the article, so there are typically many contributions per article. To give an example of how this works, PMID 18161020 is published in the journal Aids and Behavior, which is categorised under two journal categories: Behavioral Sciences (Journal category 31) and Acquired Immunodeficiency Syndrome (Journal category 86). This article includes six words/phrases (“close friend”, “customer”, “multiple sexual partners”, “risk behaviors”, “selfefficacy”, “sexual partner”) that each belong to one category set (“Group”, “Conceptual Entity”, “Finding”, “Individual Behavior”, “Mental Process”, and “Population Group”, respectively), one phrase (“alcohol use”) that belongs to two category sets (“Clinical Attribute” and “Mental or Behavioral Dysfunction”), and one word (“spouse”) that belongs to the 3 category sets noted above. It thus is recorded as containing 22 contributions, 6 x 1 x 2 = 12, plus 1 x 2 x 2 = 4, plus 1 x 3 x 2 = 6.
This multiple counting of specific ideas (i.e. words/phrases) might seem reasonable if allocation of ideas to categories was straightforward, but I didn't find it so. Some idea categories such as "Finding", "Idea or Concept" or "Conceptual Entity" are extremely vague and I didn't understand the logic of why some phrases were assigned to them.
Furthermore, the coding of ideas for each article also seemed unintuitive. The title of PMID 18161020 is Multiple Sexual Partnerships in a Sample of African-American Crack Smokers, but I could not find the terms African-American, crack, cocaine or condom anywhere in the database. There was little overlap between the ideas coded by P&B and the MeSH codes allocated by PubMed.
This led me to check the database for terms from my own subject area. One of my research topics, "handedness'", was not included in the set of ideas coded by P&B, so I was curious as to how items on this topic were coded. I checked PubMed for "handedness" in date range 2010-2016, and found over 15,000 articles. The first four research articles by US authors are shown in Table 2. All had "handedness" in the title. But PMID20943944 was categorised as "handedness right" in Idea category 124 (Organism attribute), PMID25204683 as "hand preference" in Idea category 36 (Finding), PMID20639549 as Edinburgh Handedness Inventory in Idea category 114 (Intellectual product), and PMID 23398763 had no code related to handedness. There was a 30 year difference between the idea cohorts for "hand preference" (1955) and "handedness right" (1985), but I could see no justification for these different codes.
| Article | Idea (Cohort) | Category set |
|---|---|---|
| 20639549: Handedness, heritability, neurocognition and brain asymmetry in schizophrenia | edinburgh handedness inventory (1984) | 114, Intellectual product |
| 20943944: Handedness, dexterity, and motor cortical representations | handedness right (1985) | 124, Organism attribute |
| 23398763: Handedness and corpus callosal morphology in Williams syndrome/ | - | - |
| 25204683: Evaluating handedness measures in spider monkeys | hand preference (1955) | 36, Finding |
It could be that these examples of counterintuitive codes are atypical, though I found I did not have to search to find these cases - a similar picture was found with other topics (e.g. "aphasia", "autism") - odd examples of idea coding popped up as soon as I searched PubMed for these topics. It's possible there is a coherent explanation, but it is difficult to evaluate this aspect of the study because insufficient information is given as to how the dataset was created. We know it was derived from public databases from MEDLINE and UMLS, but it remains unclear precisely how these two sources of information were combined.
Problems with the method of analysis
I also found that results depended crucially on how the dependent variable, novelty of ideas, was defined. It turns out there are various ways you can quantify this, and the choice of analysis has a big impact on the results.
P&B's main result, which I replicated, was presented thus:
Figure 1 from Packalen and Bhattacharya (2020)
I encourage readers to look carefully at this Figure to check if you can understand it. It looks simple but has some puzzling features.
The clear thing that springs out from the plot is that the seven points on the right hand end of the y-axis fall below the dotted line, which corresponds to the mean rate of NIH funding overall for this set of US-based research studies.
This result is the basis for Bhattacharya's statement "papers in 2017 were relying on ideas that were like 7 or 8 years old", echoed in the Abstract of P&B (2020) with the statement
"Papers that build on very recent ideas are NIH funded less often than are papers that build on ideas that have had a chance to mature for at least 7 y. "
There are two odd aspects of this analysis. The first is the x-axis, showing "cohort of newest idea input". Figure 1 is based on all 992,633 US-based biomedical research papers published during 2010 to 2016, but the date of publication is not considered. This makes little sense, because a paper cannot include an idea that is more recent than its publication date. Furthermore, an idea of cohort 2010 would be new if included in a paper published in 2011, but relatively old in a paper published in 2016. Such subtleties are missed in Figure 1.
Second, papers are not represented in Figure 1. As noted above, what you are seeing are "contributions". As the authors explain "We count a paper that is linked to K idea types and J research areas as K*J contributions".
Because Figure 1 looks at contributions rather than papers, it misses the key point about many papers, which is that their contribution represents a combination of ideas. Suppose an NIH-funded article reported a new treatment (cohort 2016) applied to five well-known diseases (cohort in 1950s) - this would contribute five datapoints for old ideas to Figure 1 and only one for a novel idea.
It struck me that a better way of indexing novelty of ideas in NIH-funded research would be to treat the paper as the unit of analysis, and identify the most recent idea included in a paper, measured as the smallest gap between the cohort of the idea and the year of publication. If you do that, and plot the data as in Figure 1, you get Figure 2.
Figure 2: Plot as for Figure 1, but with paper as the unit of analysis, where x-axis shows gap between idea 'cohort' and publication year of the paper featuring that idea
As you can see, the evidence now goes directly against the conclusion "Papers that build on very recent ideas are NIH funded less often than are papers that build on ideas that have had a chance to mature for at least 7 y." The points for a gap of 1 to 7 years are above the dotted line - the point for zero years is the only one slightly below the line. So, it now looks as if ideas from the past 7 years are more likely to be funded by NIH than by other funders.
It's always concerning when a modification of the analysis reverses the outcome, and I contacted the authors to ask what they made of this.
Dr Packalen gave a detailed reply, referring to the famous "garden of forking paths" and the numerous decisions that have to be made at each step of analysis that potentially would allow for thousands of analyses. He mentioned a number of factors that might affect which analysis was selected, including pressure from the academic community to obtain a specific result, and ease of interpretation of the results.
Neither of these satisfied me, because I had not found the published analysis to be easy to interpret - quite the opposite, and I was concerned that if a change in analysis could flip the finding from NIH looking weak on edge science to it looking strong on edge science, then this would merit some discussion.
Regarding the reliance on "contributions" rather than articles, Dr Packalen replied that this was a consequence of rejecting "the Great Man Theory of Science". But an analysis using the paper as the unit of analysis does not entail we believe in the Great Man Theory of Science - it just means that an article contributes to knowledge with a constellation of ideas that come together, and that to separate them is like evaluating the deliciousness of a cake from a list of its ingredients.
In a previous article, Packalen and Bhattacharya (2017) had addressed similar issues from their large database to ask which journals published most novel ideas. I pointed out that in that article, they adopted the analytic approach I used for Figure 2, i.e. focusing on the gap between publication date and cohort, and taking the shortest gap for each article. I asked why they had abandoned this approach - which seemed more logical to me - in the 2020 paper.
Dr Packalen replied saying that both the decision to base the analysis on contributions rather than papers, and to focus on absolute date of ideas, rather than novelty relative to publication year, were "forks", implying that some degree of arbitrariness is to be expected in such an analysis. Indeed, the analyses included in the reproducibility package are numbered 66, 88, 89, 90, 91 and 93, suggesting there were many more unreported analyses. Dr Packalen pointed out that some alternative analyses had been provided in supplementary material. While this is the case, these analyses all used idea cohort on the x-axis, and as far as I could see, none looked at the gap between the idea cohort and publication date (as in Figure 2).
The impression was that he was quite relaxed about accepting that there were many possible analyses and they might give different answers, implying that perhaps it was naive to imagine there could be one true analysis. I suspect this viewpoint may be normative in his field. But, as Opoku-Agyemang pointed out in this blogpost
"While it is encouraging that a growing number of journals are requiring authors to provide all data and code, the conference audience agreed that the code submitted to a journal typically represents a 'selection bias' of sorts, relative to the universe of queries actually implemented on the data. Thus, just replicating what authors provide is a frightfully low bar for transparency because we never know what non-submitted queries were run on a dataset and how these queries may have changed the outcome of the research."
Given my background in psychology, I'm not at all relaxed about wandering around the garden of forking paths and selecting your analysis after inspecting the results. We know that approach leads to non-reproducible studies and much has been written about how to overcome the resulting bias. In experimental disciplines, one solution is to adopt the Registered Reports approach, where you preregister your methods and analysis plan, and have your paper reviewed prior to results being collected.
If you are analysing an existing dataset, that's not feasible (unless you can prove you have not peeked at the data prior to developing an analysis plan). In such a situation some might recommend conducting a range of analyses and reporting them all - in the extreme case doing a multiverse analysis. However, as Lakens et al (2026) have argued, different approaches to analysis entail different auxiliary assumptions about the link between theory and empirical observations; when a multiverse analysis reveals different consequences of analytic decisions, this may help clarify what auxiliary assumptions are involved, but at the end of the day, some assumptions are theoretically justified and others are not. In a similar vein, Bishop & Hulme (2024) looked at a real life example of alternative analyses of an educational trial, and rejected the idea that all alternative analyses are equally valid. We suggested that a good way to clarify underlying theoretical assumptions is by using simulation to generate data from a model, and then testing how well different analytic approaches captured the underlying model. It would be fascinating to try this approach with the edge science data.
In my view it is not defensible to draw conclusions about the novelty of ideas in published articles by using a dependent variable that does not use the article as the unit of analysis, and that neglects to take into account the gap between publication date and idea vintage. This seems epistemically incoherent. The problem is made worse by the lack of clarity of how ideas were assigned to articles.
Having said all that, there's no doubt that having a Reproducibility Package is a huge advance on the times when no data was made public. It has made it possible for me to evaluate the dataset and the analytic approach in a way that would have been otherwise impossible, and I am grateful to the authors for making the materials available and engaging with my queries**.
However, I fear that if we combine a relaxed approach to researcher choice of forking paths with a stringent requirement for open data and code, we run the risk of giving unwarranted credibility to analytic decisions that may be non-optimal at best and positively misleading at worst. This is particularly concerning when results are used to inform major policy decisions.
Notes
*The full unedited text is appended to an earlier blogpost
**I have paraphrased the arguments used in Prof Packalen's email; I have requested his permission to reproduce the email in full, and will append it here if this is given.
As usual, I welcome on-topic comments, but seldom post them if anonymous. Comments are moderated to prevent spam and so there is typically a lag of a few days before they appear.

