Showing posts with label p-hacking. Show all posts
Showing posts with label p-hacking. Show all posts

Sunday, 16 March 2025

Book Review: Unreliable: Bias, Fraud, and the Reproducibility Crisis in Biomedical Research

by Csaba Szabo, Columbia University Press, 2025 


This is a rollicking good read, written in an informal style, and enlivened by cartoons, which works as a scholarly and accessible account of the so-called reproducibility crisis in biomedical research. 

I first became aware of this book back in February 2024 when the publisher asked me to review the draft. By happy coincidence I had just submitted an (ultimately unsuccessful) application for funding for a meeting on the closely-related topic of research fraud, and as I devoured the text, I felt guilty that I had not been aware of Szabo's work. As it turned out, there was a good reason for my ignorance: he had not previously written anything on this topic. As he explains in the Afterword, there is a personal backstory: 
The journey of a starryeyed young scientist entering the field of science, and continuing with the scientist working hard and hopefully contributing to the field over thirty years, but during all this time gradually realizing that the entire system suffers from major problems. And now that same scientist—not so young anymore, unfortunately—has written a book concluding that about 30 percent of the papers that come out every year are fake garbage and that 70 to 90 percent of the published scientific literature is not reproducible. 
That is a startling statement, but Szabo speaks with authority, as one who has always loved science, and has had a long and distinguished career in biomedical research both in the US and in Europe. It is clear that he does not want to attack science: as he points out, it is the only game in town. But he is dismayed at how the scientific method has been degraded, and he is concerned that nobody in power is taking responsibility for cleaning it up. 

What makes the book unlike any other on this topic is the detailed account of how we got into the current state, and what are the barriers to remedying the situation. Szabo spends some time explaining how hypercompetition for research grants drives the behaviour of researchers. Although it is customary to talk of "publish or perish", Szabo argues that the real crunch point for a biomedical scientist is success in obtaining external grants. In the USA, this usually means NIH funding in the form of a R01 grant, typically around $1 million over a 4-5 year period. While that sounds like a lot of money, it has to cover 50% of the salary of the principal investigator as well as other salaries and research supplies, which means it is not enough to support a research group. The success rate is around 20%, and researchers typically need to submit numerous proposals in order to survive. The institutions want their staff to obtain grants, not just so they can bathe in the reflected glory of impressive research results, but also because grants bring overheads in the form of indirect costs. So getting grants is extremely high-stakes.

It gets even more interesting when Szabo documents his experiences as a grant reviewer for NIH study sections. We might hope that the most successful proposals are the ones that are realistic, avoid hype, and carefully document the reliability of their methods. Alas, this is the opposite of what happens. Such is the pressure for novelty and impact, that anyone who proposed to replicate a prior finding would be quickly triaged out of the competition. I've noted similar tendencies in the UK context. Thus, the message researchers get from institutions is, if you want to keep your job, get a grant, and the message they get from funders is, if you want to get a grant, concentrate on making the research look exciting.

Szabo next moves on to discuss the way science is done in the lab. There are numerous factors that conspire to make findings irreproducible. Some of these are related to the inherent variability of biological systems, but some arise because of failure to adopt experimental designs that adequately control for bias. Typically, doing things meticulously takes time, and the pressure is to get results out fast. Furthermore, many studies involve a mixture of complicated methods, and the principal investigator may not understand all of them. When it comes to analysing the results, there is huge scope for adopting methods such as post-hoc outlier exclusion, p-hacking, and HARKing. All of these can be used to squeeze positive findings out of an unpromising dataset. If biomedical science is like psychology, many researchers regard such methods as normative and are unaware just how much they contribute to lack of reproducibility of published results. 

The next chapter takes a darker turn, moving to intentional fraud. A number of high-profile cases are reviewed, as well as the industrial-scale fraudulent operations run by so-called paper mills. This chapter is particularly depressing as Szabo takes us through the analogies that have been used to characterise fraud.  Initially, fraud was seen as rare, due to a few "bad apples"; later it was compared to an iceberg of fraudulent work, where we only see the tip but need to be aware that much is hidden. Szabo writes: 
In my view, even this analogy is severely misleading. If we want to stay with nature analogies, my feeling is that we are dealing with a big scientific swamp, with various swamp creatures of different sizes and shapes living in it. There are some relatively clean areas of water, too, and there are regular life forms as well. But there are an awful lot of swamp creatures who happily coexist in their natural environment, taking away food and resources from the regular life forms. In addition, a whole ecosystem built around the swamp is benefiting from it. The people who are supposed to manage the swamp, or perhaps drain it, are nowhere to be found. 
One group of people who might be expected to manage the swamp are those who publish research papers, but Szabo does not find them equal to the task, instead talking of "A broken scientific publishing system". Regular readers of this blog will be well-acquainted with the phenomenon whereby someone reports an obvious problem with a published paper, only to be ignored. Academic publishers are making efforts to screen new submissions for plagiarism and image manipulation, but there seems little appetite for cleaning up the existing body of scientific literature. Until that is done, we cannot regard it as a foundation for future work. 

Szabo is impressed by the efforts of "data sleuths", who perform post-publication peer review and report problems on the PubPeer website, but he regards this as unsustainable: and cleaning up the literature should not be a task for volunteers.  It seems that everyone wants someone to "do something" to fix the problem, but nobody takes it on. Organisations with some responsibility include universities, research institutes, publishers, editors and funders. Szabo's recommendations for change focus on funders, who have the power to deny funding to those who fail to take steps to ensure that their results are reliable. And ultimately, the money for research comes from taxpayers, and governments call the shots. 

This is a particularly difficult time to be conveying such a message. The only people who might be overjoyed to hear that a high proportion of published research is unreliable are politicians who are antagonistic to science and would like an excuse to defund it. In the USA, cuts to funding have been so fast and so deep that many are fearful that the science base may not recover. The swamp creatures may die, but so will the regular life forms. We urgently need therefore to look seriously at recommendations for changing how the system works at all levels - laboratory practice, funding, institutional integrity investigations, publishing, incentive structures - so that we can not only have confidence in scientific findings, but also defend science against attacks. 

I don't agree with all of Szabo's recommendations, but it is refreshing to have someone take a deep dive into the topic, and his ideas form a good basis for discussion. One point where I have a different approach concerns the emphasis on replications. There are many people arguing that more funding should be directed towards replicating prior studies. In the short term, that will be needed, because the way research has been done means we don't know which findings are solid. But in many areas it takes a large amount of time and money to replicate a study. The whole point of the statistical and experimental methods used in science is that they should allow us to assign a level of confidence in our findings without needing to perform an explicit replication. The problem is that we have misapplied those methods. The Registered Reports approach, where a study is evaluated by reviewers and accepted or rejected by a journal on the basis of introduction, methods and analysis plan, before any data are gathered, gets rid of the biases due to p-hacking, HARKing and publication bias, making it possible to interpret statistics sensibly. It also leads to improved methods overall, because independent reviewers offer feedback at a point when it can be helpful. As far as I know, the Registered Reports model has not been adopted by biomedicine, but could, I think, transform the field to make it more rigorous. 

Finally, I'm pleased to say that despite the initial rejection, I managed eventually to secure funding for that meeting on research fraud, which will be held in Oxford from 7th-9th April 2025, and will provide a great opportunity for discussing these issues. Registration is open for a few more days, so please consider attending (online or in person) if you'd like to take part. More details and registration form here: (turn off VPN if it does not load).

Sunday, 24 March 2024

Just make it stop! When will we say that further research isn't needed?

 

I have a lifelong interest in laterality, which is a passion that few people share. Accordingly, I am grateful to RenĂ© Westerhausen who runs the Oslo Virtual Laterality Colloquium, with monthly presentations on topics as diverse as chiral variation in snails and laterality of gesture production. 

On Friday we had a great presentation from Lottie Anstee who told us about her Masters project on handedness and musicality. There have been various studies on this topic over the years, some claiming that left-handers have superior musical skills, but samples have been small and results have been mixed. Lottie described a study with an impressive sample size (nearly 3000 children aged 10-18 years) whose musical abilities were evaluated on a detailed music assessment battery that included self-report and perceptual evaluations. The result was convincingly null, with no handedness effect on musicality. 

What happened next was what always happens in my experience when someone reports a null result. The audience made helpful suggestions for reasons why the result had not been positive and suggested modifications of the sampling, measures or analysis that might be worth trying. The measure of handedness was, as Lottie was the first to admit, very simple - perhaps a more nuanced measure would reveal an association? Should the focus be on skilled musicians rather than schoolchildren? Maybe it would be worth looking at nonlinear rather than linear associations? And even though the music assessment was pretty comprehensive, maybe it missed some key factor - amount of music instruction, or experience of specific instruments. 

After a bit of to and fro, I asked the question that always bothers me. What evidence would we need to convince us that there is really no association between musicality and handedness? The earliest study that Lottie reviewed was from 1922, so we've had over 100 years to study this topic. Shouldn't there be some kind of stop rule? This led to an interesting discussion about the impossibility of proving a negative and whether we should be using Bayes Factors, and what would be the smallest effect size of interest.  

My own view is that further investigation of this association would prove fruitless. In part, this is because I think the old literature (and to some extent the current literature!) on factors associated with handedness is at particular risk of bias, so even the messy results from a meta-analysis are likely to be over-optimistic. More than 30 years ago, I pointed out that laterality research is particularly susceptible to what we now call p-hacking - post hoc selection of cut-offs and criteria for forming subgroups, which dramatically increase the chances of finding something significant. In addition, I noted that measurement of handedness by questionnaire is simple enough to be included in a study as a "bonus factor", just in case something interesting emerges. This increases the likelihood that the literature will be affected by publication bias - the handedness data will be reported if a significant result is obtained, but otherwise can be disregarded at little cost. So I suspect that most of the exciting ideas about associations between handedness and cognitive or personality traits are built on shaky foundations, and would not replicate if tested in well-powered, preregistered studies.  But somehow, the idea that there is some kind of association remains alive, even if we have a well-designed study that gives a null result.  

Laterality is not the only area where there is no apparent stop rule. I've complained of similar trends in studies of association between genetic variants and psychological traits, for instance, where instead of abandoning an idea after a null study, researchers slightly change the methods and try again. In 2019, Lisa Feldman Barrett wrote amusingly about zombie ideas in psychology, noting that some theories are so attractive that they seem impossible to kill. I hope that as preregistration becomes more normative, we may see more null results getting published, and learn to appreciate their value. But I wonder just what it takes to get people to conclude that a research seam has been mined to the point of exhaustion. 


Monday, 4 September 2023

Polyunsaturated fatty acids and children's cognition: p-hacking and the canonisation of false facts

One of my favourite articles is a piece by Nissen et al (2016) called "Publication bias and the canonization of false facts". In it, the authors model how false information can masquerade as overwhelming evidence, if, over cycles of experimentation, positive results are more likely to be published than null ones. But their article is not just about publication bias: they go on to show how p-hacking magnifies this effect, because it leads to a false positive rate that is much higher than the nominal rate (typically .05).

I was reminded of this when looking at some literature on polyunsaturated fatty acids and children's cognition. This was a topic I'd had a passing interest in years ago when fish oil was being promoted for children with dyslexia and ADHD. I reviewed the literature back in 2008 for a talk at the British Dyslexia Association (slides here). What was striking then was that, whilst there were studies claiming positive effects of dietary supplements, they all obtained different findings. It looked suspicious to me, as if authors would keep looking in their data, and divide it up every way possible, in order to find something positive to report – in other words, p-hacking seemed rife in this field.

My interest in this area was piqued more recently simply because I was looking at articles that had been flagged up because they contained "tortured phrases". These are verbal expressions that seem to have been selected to avoid plagiarism detectors: they are often unintentionally humorous, because attempts to generate synonyms misfire. For instance, in this article by Khalid et al, published in Taylor and Francis' International Journal of Food Properties we are told: 

"Parkinson’s infection is a typical neurodegenerative sickness. The mix of hereditary and natural variables might be significant in delivering unusual protein inside explicit neuronal gatherings, prompting cell brokenness and later demise" 

And, regarding autism: 

"Chemical imbalance range problem is a term used to portray various beginning stage social correspondence issues and tedious sensorimotor practices identified with a solid hereditary part and different reasons."

The paper was interesting, though, for another reason. It contained a table summarising results from ten randomized controlled trials of polyunsaturated fatty acid supplementation in pregnant women and young children. This was not a systematic review, and it was unclear how the studies had been selected. As I documented on PubPeer,  there were errors in the descriptions of some of the studies, and the interpretation was superficial. But as I checked over the studies, I was also struck by the fact that all studies concluded with a claim of a positive finding, even when the planned analyses gave null results. But, as with the studies I'd looked at in 2008, no two studies found the same thing. All the indicators were that this field is characterised by a mixture of p-hacking and hype, which creates the impression that the benefits of dietary supplementation are well-established, when a more dispassionate look at the evidence suggests considerable scepticism is warranted.

There were three questionable research practices that were prominent. First, testing a large number of 'primary research outcomes' without any correction for multiple comparisons. Three of the papers cited by Khalid did this, and they are marked in Table 1 below with "hmm" in the main analysis column. Two of them argued against using a method such as Bonferroni correction:

"Owing to the exploratory nature of this study, we did not wish to exclude any important relationships by using stringent correction factors for multiple analyses, and we recognised the potential for a type 1 error." (Dunstan et al, 2008)

"Although multiple comparisons are inevitable in studies of this nature, the statistical corrections that are often employed to address this (e.g. Bonferroni correction) infer that multiple relationships (even if consistent and significant) detract from each other, and deal with this by adjustments that abolish any findings without extremely significant levels (P values). However, it has been validly argued that where there are consistent, repeated, coherent and biologically plausible patterns, the results ‘reinforce’ rather than detract from each other (even if P values are significant but not very large)" (Meldrum et al, 2012)
While it is correct that Bonferroni correction is overconservative with correlated outcome measures, there are other methods for protecting the analysis from inflated type I error that should be applied in such cases (Bishop, 2023).

The second practice is conducting subgroup analyses: the initial analysis finds nothing, so a way is found to divide up the sample to find a subgroup that does show the effect. There is a nice paper by Peto that explains the dangers of doing this. The third practice, looking for correlations between variables rather than main effects of intervention: with sufficient variables, it is always possible to find something 'significant' if you don't employ any correction for multiple comparisons. This inflation of false positives by correlational analysis is a well-recognised problem in the field of neuroscience (e.g. Vul et al., 2008).

Given that such practices were normative in my own field of psychology for many years, I suspect that those who adopt them here are unaware of how serious a risk they run of finding spurious positive results. For instance, if you compare two groups on ten unrelated outcome measures, then the probability that something will give you a 'significant' p-value below .05 is not 5% but 40%. (The probability that none of the 10 results is significant is .95^10, which is .6. So the probability that at least one is below .05 is 1-.6 = .4). Dividing a sample into subgroups in the hope of finding something 'significant' is another way to multiply the rate of false positive findings. 

In many fields, p-hacking is virtually impossible to detect because authors will selectively report their 'significant' findings, so the true false positive rate can't be estimated. In randomised controlled trials, the situation is a bit better, provided the study has been registered on a trial registry – this is now standard practice, precisely because it's recognised as an important way to avoid, or at least increase detection of, analytic flexibility and outcome switching. Accordingly, I catalogued, for the 10 studies reviewed by Khalid et al, how many found a significant effect of intervention on their planned, primary outcome measure, and how many focused on other results. The results are depressing. Flexible analyses are universal. Some authors emphasised the provisional nature of findings from exploratory analyses, but many did not. And my suspicion is that, even if the authors add a word of caution, those citing the work will ignore it.  


Table 1: Reporting outcomes for 10 studies cited by Khalid et al (2022)

Khalid # Register N Main result* Subgrp Correlatn Abs -ve Abs +ve
41 yes 86 NS yes no no yes
42 no 72 hmm no no no yes
43 no 420 hmm no no yes yes
44 yes 90 NS no yes yes yes
45 no 90 yes no yes NA
yes
46 yes 150 hmm no no yes yes
47 yes 175 NS no yes yes yes
48 no 107 NS yes no yes yes
49 yes 1094 NS yes no yes yes
50 no 27 yes no no yes yes

Key: Main result coded as NS (nonsignificant), yes (significant) or hmm (not significant if Bonferroni corrected); Subgrp and Correlatn coded yes or no depending on whether post hoc subgroup or correlational analyses conducted. Abs -ve coded yes if negative results reported in abstract, no if not, and NA if no negative results obtained. Abs +ve coded yes if positive results mentioned in abstract.

I don't know if the Khalid et al review will have any effect – it is so evidently flawed that I hope it will be retracted. But the problems it reveals are not just a feature of the odd rogue review: there is a systemic problem with this area of science, whereby the desire to find positive results, coupled with questionable research practices and publication bias, have led to the construction of a huge edifice of evidence based on extremely shaky foundations. The resulting waste in researcher time and funding that comes from pursuing phantom findings is a scandal that can only be addressed by researchers prioritising rigour, honesty and scholarship over fast and flashy science.

Wednesday, 29 June 2022

A proposal for data-sharing that discourages p-hacking

Open data is a great way of helping give confidence in the reproducibility of research findings. Although we are still a long way from having adequate implementation of data-sharing in psychology journals (see, for example, this commentary by Kathy Rastle, editor of Journal of Memory and Language), things are moving in the right direction, with an increasing number of journals and funders requiring sharing of data and code. But there is a downside, and I've been thinking about it this week, as we've just published a big paper on language lateralisation, where all the code and data are available on Open Science Framework. 

One problem is p-hacking. If you put a large and complex dataset in the public domain, anyone can download it and then run multiple unconstrained analyses until they find something, which is then retrospectively fitted to a plausible-sounding hypothesis. The potential to generate non-replicable false positives by such a process is extremely high - far higher than many scientists recognise. I illustrated this with a fictitious example here

Another problem is self-imposed publication bias: the researcher runs a whole set of analyses to test promising theories, but forgets about them as soon as they turn up a null result. With both of these processes in operation, data sharing becomes a poisoned chalice: instead of increasing scientific progress by encouraging novel analyses of existing data, it just means more unreliable dross is deposited in the literature. So how can we prevent this? 

In this Commentary paper, I noted several solutions. One is to require anyone accessing the data to submit a protocol which specifies the hypotheses and the analyses that will be used to test them. In effect this amounts to preregistration of secondary data analysis. This is the method used for some big epidemiological and medical databases. But it is cumbersome and also costly - you need the funding to support additional infrastructure for gatekeeping and registration. For many psychology projects, this is not going to be feasible. 

A simpler solution would be to split the data into two halves - those doing secondary data analysis only have access to part A, which allows them to do exploratory analyses, after which they can then see if any findings hold up in part B. Statistical power will be reduced by this approach, but with large datasets it may be high enough to detect effects of interest.  I wonder if it would be relatively easy to incorporate this option into Open Science Framework: i.e. someone who commits a preregistration of a secondary data analysis on the basis of exploratory analysis of half a dataset then receives a code that unlocks the second half of the data (the hold-out sample). A rough outline of how this might work is shown in Figure 1.

Figure 1: A possible flowchart for secondary data analysis on a platform such as OSF

An alternative that has been discussed by MacCoun and Perlmutter is blind analysis - "temporarily and judiciously removing data labels and altering data values to fight bias and error". The idea is that you can explore a dataset and run a planned analysis on it, but it won't be possible for the results to affect your analysis, because the data have been changed, so you won't know what is correct. A variant of this approach would be multiple datasets with shuffled data in all but one of them. The shuffling would be similar to what is done in permutation analysis - so there might be ten versions of the dataset deposited, with only one having the original unshuffled data. Those downloading the data would not know whether or not they had the correct version - only after they had decided on an analysis plan, would they be told which dataset it should be run on. 

I don't know if these methods would work, but I think they have potential for keeping people honest in secondary data analysis, while minimising bureaucracy and cost. On a platform such as Open Science Framework it is already possible to create a time-stamped preregistration of an analysis plan. I assume that within OSF there is already a log that indicates who has downloaded a dataset. So someone who wanted to do things right and just download one dataset (either a random half, or one of a set of shuffled datasets) would just need to have a mechanism that allowed them to gain access to the full, correct data after they had preregistered an analysis, similar to that outlined above.

These methods are not foolproof. Two researchers could collude - or one researcher could adopt multiple personas - so that they get to see the correct data as person X and then start a new process as person B, when they can preregister an analysis where results are already known. But my sense is that there are many honest researchers who would welcome this approach - precisely because it would keep them honest. Many of us enjoy exploring datasets, but it is all too easy to fool yourself into thinking that you've turned up something exciting when it is really just a fluke arising in the course of excessive data-mining. 

Like a lot of my blogposts, this is just a brain dump of an idea that is not fully thought through. I hope by sharing it, I will encourage people to come up with criticisms that I haven't thought of, or alternatives that might work better. Comments on the blog are moderated to prevent spam, but please do not be deterred - I will post any that are on topic. 

P.S. 5th July 2022 

Florian Naudet drew my attention to this v relevant paper: 

Baldwin, J. R., Pingault, J.-B., Schoeler, T., Sallis, H. M., & Munafò, M. R. (2022). Protecting against researcher bias in secondary data analysis: Challenges and potential solutions. European Journal of Epidemiology, 37(1), 1–10. https://doi.org/10.1007/s10654-021-00839-0

Saturday, 5 March 2016

There is a reproducibility crisis in psychology and we need to act on it


The MĂĽller-Lyer illusion: a highly reproducible effect. The central lines are the same length but the presence of the fins induces a perception that the left-hand line is longer.

The debate about whether psychological research is reproducible is getting heated. In 2015, Brian Nosek and his colleagues in the Open Science Collaboration showed that they could not replicate effects for over 50 per cent of studies published in top journals. Now we have a paper by Dan Gilbert and colleagues saying that this is misleading because Nosek’s study was flawed, and actually psychology is doing fine. More specifically: “Our analysis completely invalidates the pessimistic conclusions that many have drawn from this landmark study.” This has stimulated a set of rapid responses, mostly in the blogosphere. As Jon Sutton memorably tweeted: “I guess it's possible the paper that says the paper that says psychology is a bit shit is a bit shit is a bit shit.”
So now the folks in the media are confused and don’t know what to think.
The bulk of debate has been focused on what exactly we mean by reproducibility in statistical terms. That makes sense because many of the arguments hinge on statistics, but I think that ignores the more basic issue, which is whether psychology has a problem. My view is that we do have a problem, though psychology is no worse than many other disciplines that use inferential statistics.
In my undergraduate degree I learned about stuff that was on the one hand non-trivial and on the other hand solidly reproducible. Take for instance, various phenomena in short-term memory. Effects like the serial position effect, the phonological confusability effect, the superiority of memory for words over nonwords, are solid and robust. In perception, we have striking visual effects such as the MĂĽller-Lyer illusion, which demonstrate how our eyes can deceive us. In animal learning, the partial reinforcement effect is solid. In psycholinguistics, the difficulty adults have discriminating sound contrasts that are not distinctive in their native language is solid. In neuropsychology, the dichotic right ear advantage for verbal material is solid. In developmental psychology, it has been shown over and over again that poor readers have deficits in phonological awareness. These are just some of the numerous phenomena studied by psychologists that are reproducible in the sense that most people understand it, i.e. if I were to run an undergraduate practical class to demonstrate the effect, I’d be pretty confident that we’d get it. They are also non-trivial, in that a lay person would not just conclude that the result could have been predicted in advance.
The Reproducibility Project showed that many effects described in contemporary literature are not like that. But was it ever thus? I’d love to see the reproducibility project rerun with psychology studies reported in the literature from the 1970s – have we really got worse, or am I aware of the reproducible work just because that stuff has stood the test of time, while other work is forgotten?
My bet is that things have got worse, and I suspect there are a number of reasons for this:
1. Most of the phenomena I describe above were in areas of psychology where it was usual to report a series of experiments that demonstrated the effect and attempted to gain a better understanding of it by exploring the conditions under which it was obtained. Replication was built in to the process. That is not common in many of the areas where reproducibility of effects is contested.
2. It’s possible that all the low-hanging fruit has been plucked, and we are now focused on much smaller effects – i.e., where the signal of the effect is low in relation to background noise. That’s where statistics assumes importance. Something like the phonological confusability effect in short-term memory or a MĂĽller-Lyer illusion is so strong that it can be readily demonstrated in very small samples. Indeed, abnormal patterns of performance on short-term memory tests can be used diagnostically with individual patients. If you have a small effect, you need much bigger samples to be confident that what you are observing is signal rather than noise. Unfortunately, the field has been slow to appreciate the importance of sample size and many studies are just too underpowered to be convincing.

3. Gilbert et al raise the possibility that the effects that are observed are not just small but also more fragile, in that they can be very dependent on contextual factors. Get these wrong, and you lose the effect. Where this occurs, I think we should regard it as an opportunity, rather than a problem, because manipulating experimental conditions to discover how they influence an effect can be the key to understanding it. It can be difficult to distinguish a fragile effect from a false positive, and it is understandable that this can lead to ill-will between original researchers and those who fail to replicate their finding. But the rational response is not to dismiss the failure to replicate, but to first do adequately powered studies to demonstrate the effect and then conduct further studies to understand the boundary conditions for observing the phenomenon. To take one of the examples I used above, the link between phonological awareness and learning to read is particularly striking in English and less so in some other languages. Comparisons between languages thus provide a rich source of information for understanding how children become literate. Another of the effects, the right ear advantage in dichotic listening holds at the population level, but there are individuals for whom it is absent or reversed. Understanding this variability is part of the research process.
4. Psychology, unlike many other biomedical disciplines, involves training in statistics. In principle, this is thoroughly good thing, but in practice it can be a disaster if the psychologist is simply fixated on finding p-values less than .05 – and assumes that any effect associated with such a p-value is true. I’ve blogged about this extensively, so won’t repeat myself here, other than to say that statistical training should involve exploring simulated datasets so that the student starts to appreciate the ease with which low p-values can occur by chance when one has a large number of variables and a flexible approach to data analysis. Virtually all psychologists misunderstand p-values associated with interaction terms in analysis of variance – as I myself did until working with simulated datasets. I think in the past this was not such an issue, simply because it was not so easy to conduct statistical analyses on large datasets – one of my early papers describes how to compare regression coefficients using a pocket calculator, which at the time was an advance on other methods available! If you have to put in hours of work calculating statistics by hand, then you think hard about the analysis you need to do. Currently, you can press a few buttons on a menu and generate a vast array of numbers – which can encourage the researcher to just scan the output and highlight those where p falls below the magic threshold of .05. Those who do this are generally unaware of how problematic this is, in terms of raising the likelihood of false positive findings.
Nosek et al have demonstrated that much work in psychology is not reproducible in the everyday sense that if I try to repeat your experiment I can be confident of getting the same effect. Implicit in the critique by Gilbert et al is the notion that many studies are focused on effects that are both small and fragile, and so it is to be expected they will be hard to reproduce. They may well be right, but if so, the solution is not to deny we have a problem, but to recognise that under those circumstances there is an urgent need for our field to tackle the methodological issues of inadequate power and p-hacking, so we can distinguish genuine effects from false positives.