Showing posts with label systematic review. Show all posts
Showing posts with label systematic review. Show all posts

Saturday, 6 June 2020

Frogs or termites? Gunshot or cumulative science?


"Tell us again about Monet, Grandpa."

The tl;dr version of this post is that we're all so obsessed with doing new studies that we disregard prior literature. This is largely due to a scientific culture that gives disproportionate value to novel work. This, I argue, weakens our science.

This post has been brewing in my mind ever since I took part in a reading group about systematic reviews. We were discussing the new NIRO guidelines for systematic reviews outside the clinical trials context that are under development by Marta Topor and Jade Pickering. I'd been recommending systematic review as a useful research contribution that could be undertaken when other activities had stalled because of the pandemic. But the enthusiasm of those in the reading group seemed to wane as the session progressed. Yes, everyone agreed, the guidelines were excellent: clear and comprehensive. But it was evident that doing a proper review would not be a "quick win"; the amount of work would of course depend on the number of papers on a topic, but even for a circumscribed subject it was likely to be substantial and involve close reading of a lot of material. Was it a good use of time, people asked. I defended the importance of looking at past literature: it's concerning if we don't read scientific papers because we are all so busy writing them. To my mind, being a serious scholar means being very familiar with past work in a subject area. However, it's concerning that our reward system doesn't value that, making early-career researchers nervous about investing time in it.

The thing that prompted me to put my thoughts into words was a tweet I saw this morning by Mike Johansen (@mikejohansenmd). It seems at first to be on an unrelated topic, but I think it is another symptom of the same issue: a disregard for prior literature. Mike wrote:
Manuscripts should look like: Question: Methods: Results: Limitations: Figures/Tables: Who does these things? Things that don't matter: introduction, discussion. Who does these things?
I replied that he seemed to be recommending that we disregard the prior literature, which I think is a bad idea. I argued "One study is never enough to answer a question - important to consider how this study fits in - or if it doesn't , why."

Noah Haber (@noahhaber) jumped in at this point to say: 
I'm sympathetic (~45% convinced) to the argument that literature reviews in introductions do more harm than good. In practice, they are rarely more than cursory and uncritical, and make us beholden to ideas that have long outlived their usefulness. Space better used in methods.
But I don't think that's a good argument. I'm the first to agree that literature reviews are usually terrible: people only cite the work that confirms their position, and often do that inaccurately. You can see slides from a talk I gave on 'Why your literature review should be systematic' here. But I worry if the response to current unscholarly and biased approaches to the literature is to say that we can just disregard the literature. If you assume that the study you are doing is so important that you don't have time to read other people's studies, it is on the one hand illogical (if we all did that, who would read your studies), on the other hand disrespectful to fellow scientists, and on the most important third hand (yes, assume a mutant for now) bad for science.

Why is it bad for science? Because science seldom advances by a single study. Solid progress is made when work is cumulative. We have far more confidence in a theory that is supported by a series of experiments than by a single study, however large the effect. Indeed, we know that studies heralding a novel result often overestimate the size of effect – the "winner's curse". So to interpret your study, I want to know how far it is consistent with prior work, and if it isn't whether there might be a good reason for that.

Alas, this approach to science is discouraged by many funders and institutions: calls for research proposals are peppered with words such as "groundbreaking", "transformational", and "novel". There is a horror of doing work that is merely "cumulative". As a consequence, many researchers hop around like frogs in a lilypond, trying to land on a lilypad that is hiding buried treasure. It may sound dull, but I think we should model ourselves more on termites – we can only build an impressive edifice if we collaborate to each do our part and build on what has gone before.

Of course, the termite mound approach is a disaster if the work we try to build on is biased, poorly conducted and over-hyped. Unfortunately that is often the case, as noted by Noah. We come rather full circle here, because I think a motivation for Mike and Noah's tweets is recognition of the importance of reporting work in a way that will make it a solid foundation for a cumulative science of the future. I'm in full agreement with that. Where I disagree, though, is in how we integrate what we are doing now with what has gone before. We do need to see what we are doing as part of a cumulative, collaborative process in taking ideas forward, rather than a series of single-shot studies.

Friday, 20 July 2018

Standing on the shoulders of giants, or slithering around on jellyfish: Why reviews need to be systematic

Yesterday I had the pleasure of hearing George Davey Smith (aka @mendel_random) talk. In the course of a wide-ranging lecture he recounted his experiences with conducting a systematic review. This caught my interest, as I’d recently considered the question of literature reviews when writing about fallibility in science. George’s talk confirmed my concerns that cherry-picking of evidence can be a massive problem for many fields of science.

Together with Mark Petticrew, George had reviewed the evidence on the impact of stress and social hierarchies on coronary artery disease in non-human primates. They found 14 studies on the topic, and revealed a striking mismatch between how the literature was cited and what it actually showed. Studies in this area are of interest to those attempting to explain the well-known socioeconomic gradient in health. It’s hard to unpack this in humans, because there are so many correlated characteristics that could potentially explain the association. The primate work has been cited to support psychosocial accounts of the link; i.e., the idea that socioeconomic influences on health operate primarily through psychological and social mechanisms. Demonstration of such an impact in primates is  particularly convincing, because stress and social status can be experimentally manipulated in a way that is not feasible in humans.

The conclusion from the review was stark: ‘Overall, non-human primate studies present only limited evidence for an association between social status and coronary artery disease. Despite this, there is selective citation of individual non-human primate studies in reviews and commentaries relating to human disease aetiology’(p. e27937).

The relatively bland account in the written paper belies the stress that George and his colleague went through in doing this work. Before I tried doing one myself, I thought that a systematic review was a fairly easy and humdrum exercise. It could be if the literature were not so unruly. In practice, however, you not only have to find and synthesise the relevant evidence, but also to read and re-read papers to work out what exactly was done. Often, it’s not just a case of computing an effect size: finding the numbers that match the reported result can be challenging. One paper in the review that was particularly highly-cited in the epidemiology literature turned out to have data that were problematic: the raw data shown in scattergraphs are hard to reconcile with the adjusted means reported in a summary (see Figure below). Correspondence sent to the author apparently did not achieve a reply, let alone an explanation.

Figure 2 from Shively and Thompson (1994) Arteriosclerosis and Thrombosis Vol 14, No 5. Yellow bar added to show mean plaque areas as reported in Figure 3 (adjusted for preexperimental thigh circumference and TPC-HDL cholesterol ratio)
Even if there were no concerns about the discrepant means, the small sample size and influential outliers in this study should temper any conclusions. But those using this evidence to draw conclusions about human health focused on the ‘five-fold increase’ in coronary disease in dominant animals who became subordinate.

So what impact has the systematic review achieved? Well, the first point to note is that the authors had a great deal of difficulty getting it accepted for publication: it would be sent to reviewers who worked on stress in monkeys, and they would recommend rejection. This went on for some years: the abstract was first published in 2003, but the full paper did not appear until 2012.

The second, disappointing conclusion comes from looking at citations of the original studies reviewed by Petticrew and Davey Smith in the human health literature since their review appeared. The systematic review garnered 4 citations in the period 2013-2015 and just one during 2016-2018. The mean citations for the 14 articles covered in their meta-analysis was 2.36 for 2013-2015, and 3.00 for 2016-2018. The article that was the source of the Figure above had six citations in the human health literature in 2013-2015 and four in 2016-2018. These numbers aren’t sufficient for more than impressionistic interpretation, and I only did a superficial trawl through abstracts of citing papers, so I am not in a position to determine if all of these articles accepted the study authors’ conclusions. However, the pattern of citations fits with past experience in other fields showing that when cherry-picked facts fit a nice story, they will continue to be cited, without regard to subsequent corrections,  criticism or even retraction.

The reason why this worries me is that the stark conclusion would appear to be that we can’t trust citations of the research literature unless they are based on well-conducted systematic reviews. Iain Chalmers has been saying this for years, and in his field of clinical trials these are more common than in other disciplines. But there are still many fields where it is seen as entirely appropriate to write an introduction to a paper that only cites supportive evidence and ignores a swathe of literature that shows null or opposite results. Most postgraduates have an initial thesis chapter that reviews the literature, but it's rare, at least in psychology, to see a systematic review - perhaps because this is so time-consuming and can be soul-destroying. But if we continue to cherry-pick evidence that suits us, then we are not so much standing on the shoulders of giants as slithering around on jellyfish, and science will not progress.