Showing posts with label psychology. Show all posts
Showing posts with label psychology. Show all posts

Sunday, 12 January 2020

Should I stay or should I go? When debate with opponents should be avoided

Suppose you are invited to speak at a conference where some of the other speakers have views very different from yours. What do you do? My guess is that most academics would say you should accept. After all, we progress by evaluating claims and counterclaims, and robust debate is the lifeblood of scientific research. I'm going to argue here that there are exceptions and explain why I think responsible scientists should avoid a meeting called "Fixing Science: Practical Solutions for the Irreproducibility Crisis".

To understand this reaction, it helps to have read Merchants of Doubt: How a Handful of Scientists Obscured the Truth on Issues from Tobacco Smoke to Global Warming by Eric Conway and Naomi Oreskes (reviewed here). The synopsis from the book blurb is as follows:
The U.S. scientific community has long led the world in research on such areas as public health, environmental science, and issues affecting quality of life. Our scientists have produced landmark studies on the dangers of DDT, tobacco smoke, acid rain, and global warming. But at the same time, a small yet potent subset of this community leads the world in vehement denial of these dangers.

Merchants of Doubt tells the story of how a loose-knit group of high-level scientists and scientific advisers, with deep connections in politics and industry, ran effective campaigns to mislead the public and deny well-established scientific knowledge over four decades. Remarkably, the same individuals surface repeatedly - some of the same figures who have claimed that the science of global warming is "not settled" denied the truth of studies linking smoking to lung cancer, coal smoke to acid rain, and CFCs to the ozone hole. "Doubt is our product," wrote one tobacco executive. These 'experts' supplied it.
Uncertainty about science that threatens big businesses has been promoted by think tanks such as the Heartland Institute and Cato Institute, which receive substantial funding from those vested interests. The Fixing Science meeting has a clear overlap with those players.

The meeting first came to my attention when a mini Twitterstorm erupted after James Heathers tweeted:
Everyone's familiar with the manel, right? The all-male panel? Well, here's a whole new level for you.
Presenting: THE MANFERENCE.

I found myself wondering whether the lack of women was deliberate – maybe the organisers think you need a Y chromosome to do science – or whether it just showed they were tone-deaf to current social norms.

Once I twigged that this was organised by the National Association of Scholars, then everything fell into place. You can find the background to this organisation in their annual report here.

As several commentators have pointed out, they use the acronym NAS, which just happens to be the same as the highly respectable National Academy of Sciences: to avoid confusion here I will refer to them as NatAsSchols. The impression from their website and publications is that they are aligned with a neoliberal viewpoint and are opposed to attempts to increase diversity of race or gender in Universities.

So why is this organisation, whose mission is focused on issues such as preserving free speech and counteracting left-wing bias in Universities, running a meeting on fixing problems in science?

The NatAsSchols explains their interest in this topic as follows:
In April we published The Irreproducibility Crisis, a report on the modern scientific crisis of reproducibility—the failure of a shocking amount of scientific research to discover true results because of slipshod use of statistics, groupthink, and flawed research techniques. We launched the report at the Rayburn House Office Building in Washington, DC; it was introduced by Representative Lamar Smith, the Chairman of the House Committee on Science, Space, and Technology. This project signals our increasing commitment to address the academy’s flawed science as well as its abandonment of Western civilization and the liberal arts. We are following up The Irreproducibility Crisis with the investigation of four government agencies, including the Environmental Protection Agency. We are determined to find out just how badly irreproducible science has distorted government policy.
This makes it clear that the agenda is fundamentally a political one, designed to support the Trump administration's dismantling of environmental protections.

In 2018, Naomi Oreskes, author of Merchants of Doubt, wrote in Nature about a new 'Transparency Rule' proposed by the Environmental Protection Agency:
There is a crisis in US science, but it is not the one claimed by advocates for the rule. The crisis is the attempt to discredit scientific findings that threaten powerful corporate interests. The EPA is following a pattern that I and others have documented in regard to tobacco smoke, pollution, climate, and more. One tactic exploits the idea of scientific uncertainty to imply there is no scientific consensus. Another, seen in the latest efforts, insinuates that relevant research might be flawed. To add insult to injury, those using these tactics claim to be defending science.
February's meeting is in the same mould. The format of the meeting is cleverly constructed. The conference will be introduced and summed up by David J. Theroux (Founder, President and Chief Executive Officer of the Independent Institute and Publisher of The Independent Review) and Peter Wood, (President, NatAsSchols). Neither man has any scientific background. Theroux delighted the Heartland Institute last summer when he promoted the idea, recently publicised by Donald Trump, that wind turbines are responsible for killing numerous birds (to see this lampooned, click here)

Wood was an anthropologist who has been Provost at a small religious school, The King’s College in New York City (2005-2007), before moving to NatAsSchols. He has, as far as I can tell, no peer-reviewed publications, but he has written pieces deriding climate concerns, e.g. "the fantasies of global warming catastrophe are a kind of substitute religion, replete with a salvation doctrine, rituals of expiation, and a collection of demons to be cast out."

Another presenter is David Randall, who is Director of Research at NatAsSchols, policy advisor to the Heartland Institute and first author of the report on "The Irreproducibility of Modern Science". He is an unusual person to be authoring an authoritative report on the state of science. Web of Science turned up seven publications by him, all in politics journals, and none with any citations. His background is in history, library studies and fiction writing.

A rather puzzling choice of speaker is Richard K. Vedder, Distinguished Emeritus Professor of Economics, Ohio University and senior fellow at The Independent Institute, a think-tank founded by David Theroux. I could not find much evidence that he has shown any prior interest in science. He is Founding Director of the Center for College Affordability and Productivity in Washington, D.C and policy advisor to the Heartland Institute.

But there are also some accredited scientists on the programme, who can be divided into two camps. First, we have a set of five speakers who are aligned with NatAsSchols and/or the Heartland Institute and who have unconventional views on subjects such as climate change, pollution and gay relationships:

Elliott D. Bloom is Professor Emeritus at the Kavli Institute for Particle Astrophysics and Cosmology in the Stanford Linear Accelerator Laboratory (SLAC) and a Fellow of the American Physical Society. He has an entry on the Independent Institute website which states: "He was a member of the SLAC team with Jerome I. Friedman, Henry W. Kendall and Richard E. Taylor who received the 1990 Nobel Prize in Physics." I thought this meant he was a Nobel Laureate, but he's not listed as one. Nevertheless, he has a strong publication record in Physics. He has co-authored a presentation on "Global Warming: Fact or Fiction?", which concludes that the sun, rather than CO2 is the principal driver of climate change.

Anastasios Tsonis is Emeritus Distinguished Professor, Department of Mathematical Sciences, Atmospheric Sciences Group, University of Wisconsin Milwaukee; and Adjunct Research Scientist, Hydrologic Research Center, San Diego, California. He has worked on mathematical models of atmospheric processes and has a strong set of publications. He is a member of the academic advisory council of the Global Warming Policy Forum, a think tank founded by Nigel Lawson to combat policies designed to mitigate climate change.

Patrick J. Michaels, Senior Fellow, Competitive Enterprise Institute has a Wikipedia entry that states that "he was a senior fellow in environmental studies at the Cato Institute until Spring 2019. Until 2007 he was research professor of environmental sciences at the University of Virginia, where he had worked from 1980." Michaels also has an entry in the Website of the Heartland Institute 

Louis Anthony Cox is Professor, Department of Biostatistics and Informatics, University of Colorado and President of Cox Associates, a Denver-based applied research company specializing in quantitative risk analysis, causal modeling, advanced analytics, and operations research. He has a long list of publications. A Google search turns up an article in the Los Angeles Times which states:
The Trump administration’s reliance on industry-funded environmental specialists is again coming under fire, this time by researchers who say that Louis Anthony 'Tony' Cox Jr., who leads a key Environmental Protection Agency advisory board on air pollution, is a 'fringe' scientist and ideologue pushing policies detrimental to public health.
They refer to this paper in Science, which stated that Cox ignored consensus viewpoints on the effects of smog and particulate pollution. His work has also been criticised for its conflict with corporate interests.

Mark Regnerus, Professor, Sociology Department, University of Texas at Austin has a Wikipedia page which notes the controversy around his research on the adverse impact of a child having a parent who has been involved in a same-sex relationship. The research is funded by the Witherspoon Institute, a conservative think tank. Regnerus also contributed to an amicus brief in opposition to same-sex marriage. A sympathetic account of the controversy was published by the NatAsSchols .

Of the remaining 11 speakers, as far as I can see only one, Barry Smith (University at Buffalo), has any formal links with NatAsSchols. With a few exceptions they are psychologists/philosophers/statisticians with a specific interest in scientific reproducibility. Interest in this topic has been growing exponentially over the past 10 years, and, in general, those engaged in research in this area do so with the aim of improving scientific transparency and practice. However, they run the risk that their agenda can be weaponised to cast doubt on any particular part of scientific research that is politically or commercially inconvenient.

They will serve perfectly as foils to the five speakers whose minority views on climate/pollution/sexuality will not have to face questioning by anyone with deep expertise in those areas. I have no doubt that the reproducibility experts will have lively debate among themselves as to how the irreproducibility crisis should be fixed, in the process achieving the useful (to the organisers) goal of emphasising just what an unreliable and uncertain business science is.

Should they agree to speak at this meeting? As @briandavidearp, remarked on Twitter: "I'm wary of deliberately failing to engage/interact w/ people or organizations on the basis that they have diff moral or political commitments than me. That way balkanization and polarization lies."

That's an answer I would have agreed with a few years ago – after all, isn't that what academic life is all about? We should not just sit in our own bubble; rather we should engage respectfully with those who have different views. But this is really not about regular scientific debate. It's about weaponising the reproducibility debate to bolster the message that everything in science is uncertain – which is very convenient for those who wish to promote fringe ideas.

My view is that many of the speakers at this meeting are being played. On the one hand their presence on the programme may encourage other to agree to participate, and give false reassurance to attendees that this is a regular conference. And on the other, they will find that their arguments are scooped up by the Merchants of Doubt and used to argue that science is so uncertain that we should not accept the consensus view. We cannot be sure whether anthropogenic climate change is exaggerated, whether pollution is not really harmful, and whether gay relationships are damaging. Those who are concerned to see such ideas promoted without any debate between experts in those areas may wish to reconsider whether this meeting is really about 'Fixing Science', or whether it is rather about 'Fitting up Scientists'.

P.S.
15th January 2020
I thank Lee Jussim for engaging in the comments below. I can see that, given his understanding of the situation, it would make sense to take part in the meeting. But his understanding is different from mine, and I want to add this PS to clarify what I'm saying.

It was perhaps a mistake for me to note the neoliberal affiliations of NatAsSchols, as this appears to have given Lee the impression that my objections to the Fixing Science meeting is based on disapproval of talking to those with right-wing beliefs. I am myself on the left politically, but I agree with Brian Earp that, insofar as it is possible to do so in good faith, we should engage with those with differing views. If the meeting consisted solely of experts in philosophy of sciences/methods/metascience, with different political persuasions, I would not be warning people off – quite the contrary.

Indeed, the people I identified as belonging to the second group of speakers would seem to be exactly such a group. As I noted, I doubt that they will come to a consensus about how to 'Fix Science', but a good mix of views and perspectives is represented. No doubt Lee and others will discuss an issue that is of particular interest to me, which is how our social and cognitive biases affect how we evaluate evidence (see Bishop, 2020). Such biases are not specifically associated with left- or right-leaning politics – they affect all of us.

The problem I have with the meeting is not that the organisers are right-wing, but rather that their organisation's goals are linked to issues around higher education, and they have no credentials in science, yet they fervently advocate minority views about such topics as climate change.  Consider how bizarre it would be if, for instance, the Psychonomics Society declared that it planned to hold a meeting on 'Fixing Politics'. The NatAsScholars just doesn't have credibility in the area of scientific pratices. Alas, what they do have instead are links with funders whose vast wealth is used to attack science that threatens their vested interests. In this respect, I think the argument that 'the left-wingers are just as bad' breaks down.

But, I reiterate, the main point is not whether NatAsSchlols is left- or right-wing. It's the weird structuring of the meeting, which juxtaposes a set of experts in the 'reproducibility crisis' with a set of individuals who promote scientific views that are far from mainstream. The fact that the topics are ones that are supported by the Heartland Institute is telling, but the same strategy could in principle be used with any fringe view. Suppose you were a sceptic about evolution or vaccination, or a believer in pre-cognition. You know your arguments would not survive scrutiny by experts familiar with evidence in the area, so you don't invite those (and to be fair, it's unlikely that they'd come anyway, as there are diminishing returns in engaging with those whose minds are fixed). But what you can do is to cast doubt on all scientific evidence by inviting along those who are questioning the solidity and credibility of current scientific practices. That's what is happening here.

The general strategy has been in use for years, as documented by Conway and Oreskes, and applied to diverse topics such as tobacco dangers and acid rain, as well as climate change. The Merchants of Doubt love it when scientists themselves disagree about the nature of evidence, because it gives them a get-out-of-jail-free card.

I'm firmly of the belief we should not shove problems with science under the carpet: we need to understand the nature and extent of such problems in order to fix them. But it is a mistake to engage with those who want to exploit the presence of uncertainty to give credibility to their fringe views.

Bishop, D. V. M. (2020). The psychology of experimental psychologists: Overcoming cognitive constraints to improve research. The 45th Sir Frederic Bartlett Lecture Quarterly Journal of Experimental Psychology, 73(1), 1-19. doi:10.1177/1747021819886519

Saturday, 13 October 2018

Working memories: a brief review of Alan Baddeley's memoir

This post was prompted by Tom Hartley, who asked if I would be willing to feature an interview with Alan Baddeley on my blog.  This was excellent timing, as I'd just received a copy of Working Memories from Alan, and had planned to take it on holiday with me. It proved to be a fascinating read. Tom's interview, which you can find here, gives a taster of the content.

The book was of particular interest to me, as Alan played a big role in my career by appointing me to a post I held at the MRC Applied Psychology Unit (APU) from 1991 to 1998, and so I'm familiar with many of the characters and the ideas that he talks about in the book. His work covered a huge range of topics and collaborations, and the book, written at the age of 84, works both as a history of cognitive psychology and as a scientific autobiography.

Younger readers may be encouraged to hear that Alan's early attempts at a career were not very successful, and his career took off only after a harrowing period as a hospital porter and a schoolteacher, followed by a post at the Burden Neurological Institute, studying the effects of alcohol, where his funds were abruptly cut off because of a dispute between his boss and another senior figure. He was relieved to be offered a place at the MRC Applied Psychology Unit (APU) in Cambridge, eventually doing a doctorate there under the supervision of Conrad (whose life I wrote about here), experimenting on memory skills in sailors and postmen.

I had known that Alan's work covered a wide range of areas, but was still surprised to find just how broad his interests were. In particular, I was aware he had done work on memory in divers, but had thought that was just a minor aspect of his interests. That was quite wrong: this was Alan's main research interest over a period of years, where he did a series of studies to determine how far factors like cold, anxiety and air quality during deep dives affected reasoning and memory: questions of considerable interest to the Royal Navy among others.

After periods working at the Universities of Sussex and Stirling, Alan was appointed in 1974 as Director of the MRC APU, where he had a long and distinguished career until his formal retirement in 1995. Under his direction, the Unit flourished, pursuing a much wider range of research, with strong external links. Alan enjoyed working with others, and had collaborations around the world.  After leaving Cambridge,  he took up a research chair at the University of Bristol, before settling at the University of York, where he is currently based.

I was particularly interested in Alan's thoughts on applied versus theoretical research. The original  APU was a kind of institution that I think no longer exists: the staff were expected to apply their research skills to address questions that outside agencies, especially government, were concerned with. The earliest work was focused on topics of importance during wartime: e.g., how could vigilance be maintained by radar operators, who had the tedious task of monitoring a screen for rare but important events. Subsequently, unit staff were concerned with issues affecting efficiency of government operations during peacetime: how could postcodes be designed to be memorable? Was it safe to use a mobile phone while driving? Did lead affect children's cognitive development?  These days, applied problems are often seen as relatively pedestrian, but it is clear that if you take highly intelligent researchers with good experimental skills and pose them this kind of challenge, the work that ensues will not only answer the question, but may also lead to broader theoretical insights.

Although Alan's research included some work with neurological patients, he would definitely call himself a cognitive psychologist, and not a neuroscientist. He notes that his initial enthusiasm for functional brain imaging died down after finding that effects of interest were seldom clearcut and often failed to replicate. His own experimental approaches to evaluate aspects of memory and cognition seemed to throw more light than neuroimaging on deficits experienced by patients.

The book is strongly recommended for anyone interested in the history of psychology. As with all of Alan's writing, it is immensely readable because of his practice of writing books by dictation as he goes on long country walks: this makes for a direct and engaging style. His reflections on the 'cognitive revolution' and its impact on psychology are highly relevant for today's psychologists. As Alan says in the interview "... It's important to know where our ideas come from. It's all too tempting to think that whatever happened in the last two or three years is the cutting edge and that's all you need to know. In fact, it's probably the crest of a breaking wave and what you need to know is where that wave came from."

Saturday, 3 September 2016

Some thoughts on the Statcheck project



Yesterday, a piece in Retractionwatch covered a new study, in which results of automated statistics checks on 50,000 psychology papers are to be made public on the PubPeer website.
I had advance warning, because a study of mine had been included in what was presumably a dry run, and this led to me receiving an email on 26th August as follows:
Assuming someone had a critical comment on this paper, I duly clicked on the link, and had a moment of double-take when I read the comment.
Now, this seemed like overkill to me, and I posted a rather grumpy tweet about it. There was a bit of to and fro on Twitter with Chris Hartgerink, one of the researchers on the Statcheck project, and with the folks at Pubpeer, where I explained why I was grumpy and they defended their approach; as far as I was concerned it was not a big deal, and if nobody else found this odd, I was prepared to let it go.
But then a couple of journalists got interested, and I sent them a more detailed thoughts.
I was quoted in the Retraction Watch piece, but I thought it worth reporting my response in full here, because the quotes could be interpreted as indicating I disapprove of the Statcheck project and am defensive about errors in my work. Neither of those is true. I think the project is an interesting piece of work; my concern is solely with the way in which feedback to authors is being implemented. So here is the email I sent to journalists in full:
I am in general a strong supporter of the reproducibility movement and I agree it could be useful to document the extent to which the existing psychology literature contains statistical errors.
However, I think there are 2 problems with how this is being done in the PubPeer study.
1. The tone of the PubPeer comments will, I suspect alienate many people. As I argued on Twitter, I found it irritating to get an email saying a paper of mine had been discussed on PubPeer, only to find that this referred to a comment stating that zero errors had been found in the statistics of that paper.
I don't think we need to be told that - by all means report somewhere a list of the papers that were checked and found to be error-free, but you don't need to personally contact all the authors and clog up PubPeer with comments of this kind.
My main concern was that during an exceptionally busy period, this was just another distraction from other things. Chris Hartgerink replied that I was free to ignore the email, but that would be extremely rash because a comment on PubPeer usually means that someone has a criticism of your paper.
As someone who works on language, I also found the pragmatics of the communication non-optimal. If you write and tell someone that you've found zero errors in their paper, the implication is that this is surprising, because you don't go around stating the obvious*. And indeed, the final part of the comment basically said that your work may well have errors in it and even though they hadn't found them, we couldn't trust it.
Now at the same time as having that reaction, I appreciate this was a computer-generated message, written by non-native English speakers, that I should not take it personally, and no slur on my work was intended. And I would like to know if errors were found in my stats, and it is entirely possible that there are some, since none of us is perfect. So I don't want to over-react, but I think that if I, as someone basically sympathetic to this agenda, was irritated by the style of the communication, then the odds are this will stoke real hostility for those who are already dubious about what has been termed 'bullying' and so on by people interested in reproducibility.
2. I'll be interested to see how this pans out for people where errors are found.
My personal view is that the focus should be on errors that do change the conclusions of the paper.
I think at least a sample of these should be hand-checked so we have some idea of the error rate - I'm not sure if this has been done, but the PubPeer comment certainly gave no indication of that - it just basically said there's probably an error in your stats but we can't guarantee that there is, putting the onus on the author to then check it out.
If it's known that on 99% of occasions the automated check is accurate, then fine. If the accuracy is only 90% I'd be really unhappy about the current process as it would be leading to lots of people putting time into checking their papers on the basis of an insufficiently sensitive diagnostic. It would make the authors of the comments look frankly lazy in stirring up doubts about someone's work and then leaving them to check it out.
In epidemiology the terms sensitivity and specificity are used to refer to the accuracy of a diagnostic test. Minimally if the sensitivity and specificity of the automated stats check is known, then those figures should be provided with the automated message.

The above was written before Dalmeet drew my attention to the second paper, in which errors had been found. Here’s how I responded to that:

I hadn't seen the 2nd paper - presumably because I was not the corresponding author on that one. It's immediately apparent that the problem is that F ratios have been reported with one degree of freedom, when there should be two. In fact, it's not clear how the automated program could assign any p-value in this situation.

I'll communicate with the first author, Thalia Eley, about this, as it does need fixing for the scientific record, but, given the sample size (on which the second, missing, degree of freedom is based), the reported p-values would appear to be accurate.
  I have added a comment to this effect on the PubPeer site.


* I was thinking here of Gricean maxims, especially maxim of relation. 

Sunday, 29 May 2016

Ten serendipitous findings in psychology

The Thatcher Illusion (see below)
I'm a great fan of pre-registration of studies. It is, to my mind, the most effective safeguard against p-hacking and publication bias, the twin scourges that have led to the literature being awash with false positive findings. When combined with a more formal process, as in Registered Reports, it also allows researchers to benefit from reviewer expertise before they do the study, and to take control of the publication timeline.

But one salient objection to pre-registration comes up time and time again: if we pre-register our studies it will destroy the creative side of doing science, and turn it instead into a dull, robotic, cheerless process. We will have to anticipate what we might find, and close our eyes to what the data tell us.

Now this is both silly and untrue. For a start, there's nobody stopping anyone from doing fairly unstructured exploration, which may be the only sensible approach when entering a completely new area. The main thing in that case is to just be clear that this is what it is, and not to start applying statistical tests to the findings. If a finding has emerged from observing the data, testing it with p-values is statistically illiterate.

Nor is there any prohibition on reporting unexpected findings that emerge in the course of a study. Suppose you do a study with a pre-registered hypothesis and analysis plan, which you adhere to. Meanwhile, a most exciting, unanticipated phenomenon is observed in your experiment. If you are going down the kind of registered reports pathway used in Cortex, you report the planned experiment, and then describe the novel finding in a separate section. Hypothesis-testing and exploration are clearly delineated and no p-values are used for the latter.

In fact, with any new exciting observation, any reputable scientist would take steps to check its repeatability, to explore the conditions under which it emerges, and to attempt to develop a theory that can account for it. In effect, all that has happened is that the 'data have spoken' and suggested a new hypothesis, which could potentially be registered and evaluated in the usual way.

But would there be instances of important findings that would have been lost to history if we started using pre-registration years ago? Because I wanted examples of serendipitous findings to test this point, I asked Twitter, and lo, Twitter delivered some cracking examples. All of these predate by many years the notion of pre-registration, but note that, in all cases, having made the initial unexpected observation – either from unstructured exploratory research, or in the course of investigating something else - the researchers went on to shore up the findings with further, hypothesis-driven experiments. What they did not do is to report just the initial observation, embellished with statistics, and then move on, as if the presence of a low p-value guaranteed the truth of the result.

Here are ten phenomena well-known to psychologists that show how the combination of chance and the prepared mind can lead to important discoveries*. Where I could find one, I cite a primary source, but readers should feel free to contribute further background information.

1. Classical conditioning, Pavlov, 1902. 
The conventional account of Pavlov's discovery goes like this: He was a physiologist interested in processes of digestion and was studying the tendency of dogs to salivate when presented with food. He noted that over time, the dogs would salivate when the lab assistant entered the room, even before the food was presented, thus discovering the 'conditioned response': a response that is learned by association. A recent account is here. I was not able to find any confirmation of the serendipitous event in either Pavlov's Nobel speech, or in his Royal Society obituary, so it would be interesting to know if this described anywhere in his own writings or those of his contemporaries.

One thing that I did (serendipitously) discover from the latter source, was this intriguing detail, which makes it clear that Pavlov would never have had any truck with p-values, even if they had been in use in 1902: "He never employed mathematics even in its elementary form. He frequently said that mathematics is all very well but it confuses clear thinking almost to the same extent as statistics."

Suggested by @speech_woman @smomara1 @AglobeAgog 

2. Psychotropic drugs, 1950s 
Chance appears to have played an important role in the discovery of many psychotropic drugs in the early days of psychopharmacology. For instance, tricyclics were initially used to treat tuberculosis, when it was noticed that there was an unanticipated beneficial effect on mood. Even more striking is Hoffman's first-hand account of discovering the psychotropic effects of LSD, which he had developed as a potential circulatory stimulant. After experiencing strange sensations during a laboratory session, Hoffman returned to test the substances he had been working with, including LSD. "Even the first minimum dose of one quarter of a milligram induced a state of intoxication with very severe psychic disturbances, and this persisted for about 12 hours….This first planned experiment with LSD was a particularly terrifying experience because at the time, I had no means of knowing if I should ever return to everyday reality and be restored to a normal state of consciousness. It was only when I became aware of the gradual reinstatement of the old familiar world of reality that I was able to enjoy this greatly enhanced visionary experience".

Suggested by @ollirobinson @kealyj @neuroraf 

3. Orientation-sensitive receptive fields in visual cortex, 1959 
In his Nobel speech, David Hubel recounts how he and Torsten Wiesel were trying to plot receptive fields of visual cortex neurons using dots of light projected onto a screen, with only scant success, when they observed a cell that gave a massive response as a slide was inserted, creating a faint but sharp shadow on the retina. As he memorably put it, "over the audiomonitor, the cell went off like a machine gun". This initial observation led to a rich vein of research, but, again to quote from Hubel "It took us months to convince ourselves that we weren’t at the mercy of some optical artefact".

 Suggested by: @jpeelle @Anth_McGregor @J_Greenwood @theExtendedLuke @nikuss @sophiescott, @robustgar 

4. Right ear advantage in dichotic listening, 1961 
Doreen Kimura reported that when groups of digits were played to the two ears simultaneously, more were reported back from the right than the left ear (review here). This method was subsequently used for assessing cerebral lateralisation in neuropsychological patients, and a theory was developed that linked the right ear advantage to cerebral dominance for language. I have not been able to access a published account of the early work, but I recall being told during a visit to the Montreal Neurological Institute that it had taken time for the right ear advantage to be recognised as a real phenomenon and not a consequence of unbalanced headphones. The method of dichotic listening dated back to Broadbent or earlier, but it had originally been used to assess selective attention rather than cerebral lateralisation.

5. Phonological similarity effect in STM, 1964 
Conrad and Hull (1964) described what they termed 'acoustic confusions' when people were recalling short sequences of visually-presented letters, i.e. errors tended to involve letters that rhymed with the target letter, such as P, D, or G. In preparation for an article celebrating his 100th birthday, I recently listened to a recording of Conrad describing this early work, and explaining that when such errors were observed with auditory presentation, it was assumed they were due to mishearings. Only after further experiments did it become clear that the phenomenon arose in the course of phonological recoding in short-term memory. 

6. Hippocampal place cells, 1971 
In his 2014 Nobel lecture,  John O'Keefe describes a nice example of unconstrained exploratory research: "… we decided to record from electrodes … as the animal performed simple memory tasks and otherwise went about its daily business. I have to say that at this stage we were very catholic in our approach and expectations and were prepared to see that the cells fire to all types of situations and all types of memories. What we found instead was unexpected and very exciting. Over the course of several months of watching the animals behave while simultaneously listening to and monitoring hippocampal cell activity it became clear that there were two types of cells, the first similar to the one I had originally seen which had as its major correlate some non-specific higher-order aspect of movements, and the second a much more silent type which only sprang into activity at irregular intervals and whose correlate was much more difficult to identify. Looking back at the notes from this period it is clear that there were hints that the animal’s location was important but it was only on a particular day when we were recording from a very clear well isolated cell with a clear correlate that it dawned on me that these cells weren’t particularly interested in what the animal was doing or why it was doing it but rather they were interested in where it was in the environment at the time. The cells were coding for the animal’s location!" Needless to say, once the hypothesis of place cells had been formulated, O'Keefe and colleagues went on to test and develop it in a series of rigorous experiments.

7. McGurk effect, 1976 
In a famous paper, McGurk and McDonald reported a dramatic illusion: when watching a talking head, in which repeated utterances of the syllable [ba] are dubbed on to lip movements for [ga], normal adults report hearing [da]. Those who recommended this example to me mentioned that the mismatching of lips and voices arose through a dubbing error, and there was even the idea that a technician was disciplined for mixing up the tapes, but I've not found a source for that story. I noted with interest that the Nature paper reporting the findings does not contain a single p-value.
 
Suggested by: @criener @neuroconscience @DrMattDavis 

8. Thatcher illusion, 1980 
Peter Thompson kindly sent me an account of his discovery of the Thatcher Illusion (downloadable from here, p. 921). His goal had been to illustrate how spatial frequency information is used in vision, entailing that viewing the same image close up and at a distance will give very different percepts if low spatial frequencies are manipulated. He decided to illustrate this with pictures of Margaret Thatcher, one of which he doctored to invert the eyes and mouth, creating an impressively hideous image. He went to get sellotape to fix the material in place, but noticed that when he returned, approaching the table from the other side, the doctored images were no longer hideous when inverted. Had he had sellotape to hand, we might never have discovered this wonderful illusion.

Suggested by @J_Greenwood 

9. Repetition blindness, 1987 
Repetition blindness, described here by Nancy Kanwisher, is the phenomenon whereby people have difficulty detecting repeated words that are presented using rapid serial visual presentation (RSVP) - even when the two occurrences are nonconsecutive and differ in case. I could not find a clear account of the history of the discovery, but it seems that researchers investigating a different problem thought that some stimuli were failing to appear, and then realised these were the repeated ones.

Suggested by @PaulEDux 

10. Mirror neurons, 1992 
Giacomo Rizzolatti and colleagues were recording from cells in the macaque premotor cortex that responded when the animal reached for food, or bit a peanut. To their surprise, they noticed when testing the animals, the same cell that responded when the monkey picked up a peanut also responded when the experimenter did so (see here for summary). Ultimately, they dubbed these cells 'mirror neurons' because they responded both to the animal's own actions and when the animal observed another performing a similar action. The story that mirror neurons were first identified when they started responding during a coffee break as Rizzolatti picked up his espresso appear to be apocryphal.

Suggested by: @brain_apps @neuroraf @ArranReader @seriousstats @jameskilner @RRocheNeuro 

 *I picked ones that I deemed the clearest and best-known examples. Many thanks to all the people who suggested others.

Saturday, 5 March 2016

There is a reproducibility crisis in psychology and we need to act on it


The Müller-Lyer illusion: a highly reproducible effect. The central lines are the same length but the presence of the fins induces a perception that the left-hand line is longer.

The debate about whether psychological research is reproducible is getting heated. In 2015, Brian Nosek and his colleagues in the Open Science Collaboration showed that they could not replicate effects for over 50 per cent of studies published in top journals. Now we have a paper by Dan Gilbert and colleagues saying that this is misleading because Nosek’s study was flawed, and actually psychology is doing fine. More specifically: “Our analysis completely invalidates the pessimistic conclusions that many have drawn from this landmark study.” This has stimulated a set of rapid responses, mostly in the blogosphere. As Jon Sutton memorably tweeted: “I guess it's possible the paper that says the paper that says psychology is a bit shit is a bit shit is a bit shit.”
So now the folks in the media are confused and don’t know what to think.
The bulk of debate has been focused on what exactly we mean by reproducibility in statistical terms. That makes sense because many of the arguments hinge on statistics, but I think that ignores the more basic issue, which is whether psychology has a problem. My view is that we do have a problem, though psychology is no worse than many other disciplines that use inferential statistics.
In my undergraduate degree I learned about stuff that was on the one hand non-trivial and on the other hand solidly reproducible. Take for instance, various phenomena in short-term memory. Effects like the serial position effect, the phonological confusability effect, the superiority of memory for words over nonwords, are solid and robust. In perception, we have striking visual effects such as the Müller-Lyer illusion, which demonstrate how our eyes can deceive us. In animal learning, the partial reinforcement effect is solid. In psycholinguistics, the difficulty adults have discriminating sound contrasts that are not distinctive in their native language is solid. In neuropsychology, the dichotic right ear advantage for verbal material is solid. In developmental psychology, it has been shown over and over again that poor readers have deficits in phonological awareness. These are just some of the numerous phenomena studied by psychologists that are reproducible in the sense that most people understand it, i.e. if I were to run an undergraduate practical class to demonstrate the effect, I’d be pretty confident that we’d get it. They are also non-trivial, in that a lay person would not just conclude that the result could have been predicted in advance.
The Reproducibility Project showed that many effects described in contemporary literature are not like that. But was it ever thus? I’d love to see the reproducibility project rerun with psychology studies reported in the literature from the 1970s – have we really got worse, or am I aware of the reproducible work just because that stuff has stood the test of time, while other work is forgotten?
My bet is that things have got worse, and I suspect there are a number of reasons for this:
1. Most of the phenomena I describe above were in areas of psychology where it was usual to report a series of experiments that demonstrated the effect and attempted to gain a better understanding of it by exploring the conditions under which it was obtained. Replication was built in to the process. That is not common in many of the areas where reproducibility of effects is contested.
2. It’s possible that all the low-hanging fruit has been plucked, and we are now focused on much smaller effects – i.e., where the signal of the effect is low in relation to background noise. That’s where statistics assumes importance. Something like the phonological confusability effect in short-term memory or a Müller-Lyer illusion is so strong that it can be readily demonstrated in very small samples. Indeed, abnormal patterns of performance on short-term memory tests can be used diagnostically with individual patients. If you have a small effect, you need much bigger samples to be confident that what you are observing is signal rather than noise. Unfortunately, the field has been slow to appreciate the importance of sample size and many studies are just too underpowered to be convincing.

3. Gilbert et al raise the possibility that the effects that are observed are not just small but also more fragile, in that they can be very dependent on contextual factors. Get these wrong, and you lose the effect. Where this occurs, I think we should regard it as an opportunity, rather than a problem, because manipulating experimental conditions to discover how they influence an effect can be the key to understanding it. It can be difficult to distinguish a fragile effect from a false positive, and it is understandable that this can lead to ill-will between original researchers and those who fail to replicate their finding. But the rational response is not to dismiss the failure to replicate, but to first do adequately powered studies to demonstrate the effect and then conduct further studies to understand the boundary conditions for observing the phenomenon. To take one of the examples I used above, the link between phonological awareness and learning to read is particularly striking in English and less so in some other languages. Comparisons between languages thus provide a rich source of information for understanding how children become literate. Another of the effects, the right ear advantage in dichotic listening holds at the population level, but there are individuals for whom it is absent or reversed. Understanding this variability is part of the research process.
4. Psychology, unlike many other biomedical disciplines, involves training in statistics. In principle, this is thoroughly good thing, but in practice it can be a disaster if the psychologist is simply fixated on finding p-values less than .05 – and assumes that any effect associated with such a p-value is true. I’ve blogged about this extensively, so won’t repeat myself here, other than to say that statistical training should involve exploring simulated datasets so that the student starts to appreciate the ease with which low p-values can occur by chance when one has a large number of variables and a flexible approach to data analysis. Virtually all psychologists misunderstand p-values associated with interaction terms in analysis of variance – as I myself did until working with simulated datasets. I think in the past this was not such an issue, simply because it was not so easy to conduct statistical analyses on large datasets – one of my early papers describes how to compare regression coefficients using a pocket calculator, which at the time was an advance on other methods available! If you have to put in hours of work calculating statistics by hand, then you think hard about the analysis you need to do. Currently, you can press a few buttons on a menu and generate a vast array of numbers – which can encourage the researcher to just scan the output and highlight those where p falls below the magic threshold of .05. Those who do this are generally unaware of how problematic this is, in terms of raising the likelihood of false positive findings.
Nosek et al have demonstrated that much work in psychology is not reproducible in the everyday sense that if I try to repeat your experiment I can be confident of getting the same effect. Implicit in the critique by Gilbert et al is the notion that many studies are focused on effects that are both small and fragile, and so it is to be expected they will be hard to reproduce. They may well be right, but if so, the solution is not to deny we have a problem, but to recognise that under those circumstances there is an urgent need for our field to tackle the methodological issues of inadequate power and p-hacking, so we can distinguish genuine effects from false positives.