Friday, 20 April 2012

Getting genetic effect sizes in perspective


My research focuses on neurodevelopmental disorders - specific language impairment, dyslexia, and autism in particular. For all of these there is evidence of genetic influence. But the research papers reporting relevant results are often incomprehensible to people who aren’t geneticists (and sometimes to those who are).  This leaves us ignorant of what has really been found, and subject to serious misunderstandings.
Just as preamble, evidence for genetic influences on behaviour comes in two kinds. The first approach, sometimes referred to as genetic epidemiology or behaviour genetics allows us to infer how far genes are involved in causing individual differences by studying similarities between people who have different kinds of genetic relationship. The mainstay of this field is the twin study. The logic of twin studies is pretty simple, but the methods currently used to analyse twin data are complex. The twin method is far from perfect, but it has proved useful in helping us identify which conditions are worth investigating using the second approach, molecular genetics.
Molecular genetics involves finding segments of DNA that are correlated with a behavioural (or other phenotypic) measure. It involves laboratory work analysing biological samples of people who’ve been assessed on relevant measures. So if we’re interested in, say, dyslexia, we can either look for DNA variants that predict a person’s reading ability - a quantitative approach - or we can look for DNA variants that are more common in people who have dyslexia. There’s a range of methods that can be used, depending on whether the data come from families - in which case the relationship between individuals can be taken into account - or whether we just have a sample of unrelated people who vary on the behaviour of interest, in this case reading ability.
The really big problem comes from a tendency in molecular genetics to focus just on p-values when reporting findings. This is understandable: the field of molecular genetics has been plagued by chance findings. This is because there’s vast amounts of DNA that can be analysed, and if you look at enough things, then the odd result will pop up as showing a group difference just by chance. (See this blogpost for further explanation). The p-value indicates whether an association between a DNA variant and a behavioural measure is a solid finding that is likely to replicate in another sample.
But a p-value depends on two things: (a) the strength of association between DNA and behaviour (effect size) and (b) the sample size. Psychologists, many of whom are interested in genetic variants linked to behaviour, are mostly used to working with samples that number in the tens rather than hundreds or thousands. It’s easy, therefore, to fall into the trap of assuming that a very low p-value means we have a large effect size, because that’s usually the case in the kind of studies we’re used to. Misunderstanding can arise if effect sizes are not reported in a paper.
Suppose we have a genetic locus with two alleles, a and A, and a geneticist contrasts people with an aa genotype vs those with aA or AA (who are grouped together). We read a molecular genetics paper that reports an association between these genotypes and reading ability with p-value of .001. Furthermore, we see there are other studies in the literature reporting similar associations, so this seems a robust finding. You could be forgiven for concluding that the geneticists have found a “dyslexia gene”, or at least a strong association with a DNA variant that will be useful in screening and diagnosis. And, if you are a psychologist, you might be tempted to do further studies contrasting people with aa vs aA/AA genotypes on behavioural or neurobiological measures that are relevant for reading ability.
However, this enthusiasm is likely to evaporate if you consider effect sizes. There is a nice little function in R, compute.es, that allows you to compute effect size easily if you know a p-value and a sample size. The table below shows:
  •  effect sizes (Cohen’s d, which gives mean difference between groups in z-score units)
  •  average for each group for reading scores scaled so the mean for the aA/AA group is 100 with SD of 15
Results are shown for various sample sizes with equal numbers of aa vs aA/AA and either p =.001 or p = .0001. (See reference manual for the R function for relevant formulae, which are also applicable in cases of unequal sample size). For those unfamiliar with this area, a child would not normally be flagged up as having reading problems unless a score on a test scaled this way was 85 or less (i.e., 1 SD below the mean).

Table 1: Effect sizes (Cohen’s d) and group means derived from p-value and sample size (N)










When you have the kind of sample size that experimental or clinical psychologists often work with, with 25 participants per group, a p of .001 is indicative of a big effect, with a mean difference between groups of almost one SD. However, you may be surprised at how small the effect size is when you have a large sample. If you have a sample of 3000 or so, then a difference of just 1-2 points (or .08 SD) will give you p < .001. Most molecular genetic studies have large sample sizes. Geneticists in this area have learned that they have to have large samples, because they are looking for small effects!
It would be quite wrong to suggest that only large effect sizes are interesting. Small but replicable effects can be of great importance in helping us understand causes of disorders, because if we find relevant genes we can study their mode of action (Scerri & Schulte-Korne, 2010). But, as far as the illustrative data in Table 1 are concerned, few psychologists would regard the reading deficit associated at p of .001 or .0001 with genotype aa as of clinical significance, once the sample size exceeds 1000 per group.
Genotype aa may be a risk factor for dyslexia, but only in conjunction with other risks. On its own it doesn’t cause dyslexia.  And the notion, propagated by some commercial genetics testing companies, that you could use a single DNA variant with this magnitude of effect to predict a person’s risk of dyslexia, is highly misleading.


Further reading
Flint, J., Greenspan, R. J., & Kendler, K. S. (2010). How Genes Influence Behavior: Oxford University Press.
Scerri, T., & Schulte-Körne, G. (2009). Genetics of developmental dyslexia European Child & Adolescent Psychiatry, 19 (3), 179-197 DOI: 10.1007/s00787-009-0081-0

If you are interested in analysing twin data, you can see my blog on Twin Methods in OpenMx, which illustrates a structural equation modelling approach in R with simulated data.

Update 21/4/12: Thanks to Tom Scerri for pointing out my original wording talked of "two versions of an allele", which has now been corrected to "a genetic locus with two alleles"; as Tom noted: an allele is an allele, you can't have two versions of it.
Tom also noted that in the table, I had taken the aA/AA genotype as the reference group for standardisation to mean 100, SD 15. A more realistic simulation would take the whole population with all three genotypes as the reference group, in which case the effect size would result from the aA/AA group having a mean above 100, while the aa group would have mean below 100. This would entail that, relative to the grand population average, the averages for aa would be higher than shown here, so that the number with clinically significant deficits will be even smaller.
I hope in future to illustrate these points by computing effect sizes for published molecular genetic studies reporting links with cognitive phenotypes.

Thursday, 12 April 2012

The ultimate email auto-response

Prague Humor
Photo credit: szeke

Andy Field (who actually gave me a v. helpful response to a query via Twitter...) demonstrates how it should be done:

12 April 2012 06:34
This is an automatic reply.

I'm on study leave writing 'Discovering Statistics Using SPSS 4'. This essentially means that I'm locked in a mental Dungeon for the next 6 months in which intrusions from the outside world are like needles lancing my brain. They hurt, and hence I'm going to ignore them. If you really need to get hold of me then you should write a letter and insert it into the stale bread that they push through my cell door every morning. Or you can follow my demented ramblings (or 'progress' as some people call it) on Facebook and Twitter.

I will start to emerge back into reality (some might argue I was never in it) sometime in April 2012, at which point please resend your email if you still require a response.

Monday, 9 April 2012

BBC's 'extensive coverage' of the NHS bill

Last month, there was a remarkable disconnect between what was being reported on BBC News outlets and what was concerning many members of the public on social media. The Health and Social Care Bill was passed by Parliament on 21st March, despite massive objections from many of those working in the NHS, and those members of the general public who were aware of the bill. Evidence of this concern was apparent from the fact that a petition with 486,000 signatures was presented to the Lords by Lord David Owen on 19th March, supporting his view that consideration of the Bill should be deferred until after the Risk Register had been published. There had also been a rally on 7th March attended by thousands of NHS workers. During the month of March, when there was still an opportunity of killing the bill if the Liberal Democrats had come out against it, there appeared to be very little coverage of it by the BBC. Only after the Bill had been passed, did the BBC seem willing to run it as a news item.
There has been a fair bit of commentary on the lack of coverage, with some suggesting there may have been a deliberate conspiracy to keep quiet because of political pressures and/or vested interests of BBC executives in private health providers.
http://salt-mine.net/blog/2012/03/23/complaint-bbcqt-no-question-nhs-bill/ 
http://storify.com/isobelweinberg/bbc-coverage-of-the-nhs-bill 
http://caiwingfield.com/cms/2012/03/police-suppression-of-peaceful-pro-nhs-protest-march-17th-2012/
http://socialinvestigations.blogspot.co.uk/2012/03/lord-patten-of-barnes-bridgepoint-and.html
Like many people, I submitted a complaint to the BBC about their lack of coverage of this important topic. I received a prompt reply as follows:
Dear Dorothy 
Thank you for contacting us regarding BBC News coverage. We understand you believe BBC News did not sufficiently report on the opposition to the Health and Social Care Bill. BBC News has reported extensively on the opposition to the Health and Social Care Bill across our news programmes and bulletins since the Bill was originally proposed. We have reported on the health, political and business dimensions of the debate during our flagship news programmes and news bulletins and have heard from politicians, NHS workers, public sector workers and members of the public alike, as well as from supporters of the bill. There have been numerous protests and demonstrations held in opposition to the Government’s proposals. Such shows of opposition have been varied in size and were spread across the different stages of the bill’s formation. We believe we have accurately and fairly reflected the nature of this opposition in our news coverage. While you were unhappy about the level of coverage given to this, the political opposition to the Bill culminated in the House of Commons emergency debate on 20 March. Accordingly, the Commons debate featured heavily in our news coverage on the day and was the lead story during our main news bulletins. The Health and Social Care bill has been one of the biggest UK stories over the past few months and we believe we have afforded it the appropriate level coverage in a fair and impartial manner, allowing viewers and listeners to make up their own minds on the matter at hand
The phrase ‘extensive coverage’ did not reflect my impressions. I am not glued to the BBC, but I am a regular listener to the Radio 4 Today Programme and I was not aware of the NHS bill receiving any coverage at all. I therefore submitted a follow-up complaint asking if they could please give me details of specific programmes when the BBC had covered the NHS bill during the month of March. Again, they replied promptly and courteously. And here is what they said (with relevant sections from the websites in blue):
Dear Dr Bishop
We understand that you would like details of when opposition to the Health and Social Care Bill was covered by the BBC. Opposition to the bill has been covered on various programmes across the BBC, for example; Newsnight, The Daily Politics, Today and BBC News Online. Opposition was also covered on 'Newsnight' during a report on 9th March which looked at Liberal Democrat activist's plans to derail the bill and on the 13th March during a discussion on the future of the welfare system. A report on the 'The Daily Politics' broadcast on 13th March (at 14:09) highlighted opposition from Labour as well as the Royal College of GPs. Diane Abbott said health professionals were still opposed to the Health and Social Care Bill, which could be days away from becoming law. She said a future Labour government would overturn the act and "unpick the worst of the damage". Liberal Democrat spokesman Lord Clement-Jones said the bill was "going to get more acceptance"
More coverage was given on the 'Today' programme on the following dates:  
Sat 10th March, 0712-0714 Liberal Democrat activists will decide this morning whether Nick Clegg will face a vote on the health bill. The BBC's Robin Bryant explains why that could be bad news for the government. 
Weds 21st March, 0840-0853 The government's controversial plans to change the NHS have passed their final hurdle in Parliament after 14 months of opposition and changes in both houses. Professor Chris Ham, chief executive of the health think-tank The King's Fund, looks at what we left with now.
In addition examples of coverage on BBC News Online include: 


So, if I have understood this right,during March, the Today Programme covered the story once, in an early two-minute slot, before the Bill was passed. Other items that morning included 4 minutes on a French theme park based on Napoleon, 6 minutes on international bagpipe day and 8 minutes on Jubilee celebrations.
 



free counters

Tuesday, 3 April 2012

Phonics screening: sense and sensibility

There’s been a lot written about the new phonics test that is being introduced in UK schools in June. Michael Rosen cogently put the arguments against it on his blog this morning. A major concern is that the test involves asking children to read a list of items, and takes no account of whether they understand them. Indeed, the list includes nonwords (i.e. pronounceable letter strings, such as "doop" or "barg") as well as meaningful words. So children will be “barking at print” - a very different skill from reading for meaning.

I can absolutely see where Rosen is coming from, but he’s missing a key point. You can’t read for meaning if you can’t decode the words. It’s possible to learn some words by rote, even if you don’t know how letters and sounds go together, but in order to have a strategy for decoding novel words, you need the phonics skills. Sure, English is an irritatingly irregular language, so phonics doesn’t always give you the right answer, but without phonics, you have no strategy for approaching an unfamiliar word.
Back in 1990, Hoover and Gough wrote an influential paper in 1990 called “The Simple View of Reading”. This is clearly explained in this series of slides by Morag Stuart from the Institute of Education. It boils down to saying that in order to be an effective reader you need two things: the ability to decode words, and the ability to understand the language in a text. Some children can say the words but don’t understand what they’ve read. These are the ones Michael Rosen is worried about. They won’t be detected by a nonword reading test. They are all-too-often missed by teachers who don’t realise they are having problems because when asked to read aloud, they do fine. There’s a fair bit of research on these so-called “poor comprehenders”, and how best to help them (some of which is reviewed here). But there are other children with the opposite pattern: good language understanding but difficulties in decoding: this corresponds to classic dyslexia. There are decades of research showing that one of the most effective ways of identifying these children is to assess their ability to read novel letter sequences that they haven’t encountered before - nonwords. Nonword reading ability has also been shown to predict which children are at risk for later reading failure.  It's useful precisely because it tests children's ability to attack unfamiliar material, rather than testing what they have already learned. It's a bit like a doctor giving someone a stress test on a treadmill. They may never encounter a treadmill in everyday life, but by observing how they cope with it, the doctor can tell whether they are at risk of cardiovascular problems.

Some children don’t need explicit teaching of phonics - they pick it up spontaneously through exposure to print. But others just don’t get it unless it is made explicit. I’m coming at this as someone who sees children who just don’t get past first base in learning to read, and who fall increasingly far behind if their difficulties aren’t identified. A nonword reading test around age 6 to 7 years will help identify those children who could benefit from extra support in the classroom.
So that’s the rationale, and it is well-grounded in a great deal of reading research. But is there a downside? Potentially, there are numerous risks. It would be catastrophic if teachers got the message from this exercise that reading instruction should involve training children to read lists of words, or worse still, nonwords. Unfortunately, testing in schools is increasingly conflated with evaluation of the school, and so teaching-to-the-test is routinely done. The language comprehension side of reading is hugely important, and shouldn't be neglected. Developing children’s oral language skills is an important component of making children literate. It is also important for children to be read to, and to learn that books are a source of pleasure.
Another concern is children being identified at an early age as failing. The cutoff that is used is crucial, and there are concerns that the bar may be set too high.  Children at real risk are those who bomb on nonword reading, not those who are just a bit below average.
The impact on children’s self-perception is also key. There is already evidence that some primary school children are unduly stressed by SATS. There’s nothing more likely to put a child off reading than being given a test that they don’t understand and being told they’ve failed it. When I was at school, we had the 11+ examination that divided children into those who went to grammar school and those who didn’t. I had friends whose parents promised them a bicycle if they passed - even though there was precious little practice that you could do for the 11+, which was designed to test skills that had not been explicitly taught. Schoolfriends who failed were left with a chip on their shoulder for years. I’d hope that this reading screen is introduced in a more sensitive manner, but the onus is on parents, teachers and the media to ensure this happens. This screening test should serve as a simple diagnostic that will allow teachers to identify those children whose weak letter-sound-knowledge means that they could benefit from extra support. It should not be used to evaluate schools, make children feel they are failures, worry their parents, or support a sterile phonics-only approach to reading.

References
Connor, M. J. (2003). Pupil stress and standard assessment tasks (SATs) An update. Emotional and Behavioural Difficulties, 8(2), 101-107. doi: 10.1080/13632750300507010
Hoover, W. A., & Gough, P. B. (1990). The simple view of reading. Reading and Writing, 2, 127-160.
Nation, K., & Angell, P. (2006). Learning to read and learning to comprehend. London Review of Education, 4(1), 77–87. doi: 10.1080/13603110600574538
Rack, J. P., Snowling, M. J., & Olson, R. K. (1992). The nonword reading deficit in developmental dyslexia. Reading Research Quarterly, 27, 29-53.
Snowling, M., & Hulme, C. (2012). Interventions for children's language and literacy difficulties International Journal of Language & Communication Disorders, 47 (1), 27-34 DOI: 10.1111/j.1460-6984.2011.00081.x

 
free counters

Wednesday, 28 March 2012

C’mon sisters! Speak out!


When I give a talk, I like to allow time for questions. It’s not just a matter of politeness to the audience, though that is a factor. I find it helps me gauge how the talk has gone down: what points have people picked up on, are there things they didn’t get, and are there things I didn’t get? Quite often a question coming from left field gives me good ideas. Sometimes I’m challenged and that’s good too, as it helps me either improve my arguments or revise them. But here’s the thing. After virtually every talk I give there’s a small queue of people who want to ask me a private question. Typically they’ll say, “I didn’t like to ask you this in the question period, but…”, or “This probably isn’t a very sensible thing to ask, but…”. And the thing I’ve noticed is that they are almost always women. And very often I find myself saying, “I wish you’d asked that question in public, because I think there are lots of people in the audience who’d have been interested in what you have to say.”

I’m not an expert in gender studies or feminism, and most of my information about research on gender differences comes from Virginia Valian’s scholarly review, Why So Slow. Valian reviews studies confirming that women are less likely than men to speak out in question sessions in seminars. I have to say my experience in the field of psychology is rather different, and I'm pleased to work in a department where women’s voices are as likely to be heard as men’s. But there’s no doubt that this is not the norm for many disciplines, and I've attended conferences, and given talks, where 90% of questions come from men, even when they are a minority of the audience.

So what’s the explanation? Valian recounts personal experiences as well as research evidence that women are at risk of being ignored if they attempt to speak out, and so they learn to keep quiet. But, while I'm sure there is truth in that, I find myself irritated by what I see as a kind of passivity in my fellow women. It seems too easy to lay the blame at the feet of nasty men who treat you as if you are invisible. A deeper problem seems to be that women have been socially conditioned to be nervous of putting their heads above the parapet. It is really much easier to sit quietly in an audience and think your private thoughts than to share those thoughts with the world, because the world may judge you and find you lacking. If you ask women why they didn’t speak up in a seminar, they’ll often say that they didn’t think their question was important enough, or that it might have been wrong-headed. They want to live life safely and not draw attention to themselves. This affects participation in discussion and debate at all stages of academic life - see this description of anxiety about participating in student classes. Of course, this doesn’t only affect women, nor does it affect all women. But it affects enough women to create an imbalance in who gets heard.
We do need to change this. Verbal exchanges after lectures and seminars are an important part of academic life, and women need to participate fully. There’s no point in encouraging men to listen to women’s voices if the women never speak up. If you are one of those silent women, I urge you to make an effort to overcome your bashfulness. You’ll find it less terrifying than you imagine, and it gets easier with practice. Don’t ask questions just for the sake of it, but when a speaker sparks off an interesting thought, a challenging question, or just a need for clarification, speak out. We need to change the culture here so that the next generation of women feel at ease in engaging in verbal academic debate.


Tuesday, 20 March 2012

The REF: a monster that sucks time and money from academic institutions


I’ve long had a pretty cynical attitude towards the periodic exercises for rating research activities of UK higher education institutions. The problem is cost-effectiveness. Institutions put forward detailed submissions in which the best research outputs of their academics are documented. A panel of top academics then considers these, and central funding is awarded according to the ratings. This takes up massive amounts of time of those writing and reviewing the submissions.The end result is a rank ordering of institutions that seldom contains any surprises. Pretty much the same ordering could be obtained by, for instance, taking a panel of top academics in a given field and sitting them down in a room to vote. This is a point made many years ago by Colin Blakemore, talking about the REF’s predecessor, the RAE.
For the REF 2014, the rules have now changed, so that we don’t only have to document what research we’ve published: we also have to demonstrate its impact. Impact has a very specific definition. It has to originate from a particular academic paper and you have to be able to quantify its effect in the wider world. This poses a new challenge to those preparing REF submissions, a challenge that many institutions are taking very seriously. All over the UK, meetings are being convened to discuss impact statements. Here in Oxford, we’ve already had several long meetings of senior professors devoted just to this issue. Then this week I saw an advertisement from UCL that goes a step further. They are looking for three editorial consultants on a salary of £32,055 - £38,744 per annum to work on their REF impact statements.
This induced in me a sense of despair. Why, you may ask? After all, academics are hard-pressed and this is a way of taking some of the burden from them, while ensuring that their work is presented in the best possible light. My problem with this is that it exemplifies a shift in priorities from substance to presentation. Funds that could have been used to support the university’s core functions of teaching or research go towards PR. And the REF, an exercise that is supposed to enhance the UK’s research, ends up leaching money as well as time from the system.
Here’s a suggestion. Those who are on REF panels should do their own private rankings of higher education institutions and put them in a sealed envelope now. After the REF results are announced, they can compare the outcome with their predictions. Then we will be able to see whether the huge amounts of time and money spent on this exercise have been worthwhile.

Sunday, 11 March 2012

A letter to Nick Clegg from an ex Liberal Democrat


Dear Nick Clegg
Yesterday I tweeted to see if anyone following #LDConf could tell me what the party stood for. It was a serious question, but so far no response.

I've voted Lib Dem for many years. My leanings are to the left but Labour’s perpetual internal wrangling was off-putting and the Iraq War an appalling mistake. I liked a lot of the LD policies and a few years ago joined the party. I’ve always had enormous respect for Evan Harris, who was my local MP until he was ousted at the last election.

I was initially sympathetic to the idea of working in Coalition and anticipated that the Lib Dems would act as a moderating force on the excesses of the Tories. In the early months, we saw precious little of that. I started to worry when tuition fees came and went with little sign of any Lib Dem protest. Changes to disability benefits were the next thing. I’d hoped Vince Cable would be able to tackle regulation of the banks and he’s shown willing but appears ultimately toothless. While all this was going on, I found myself wondering whether there would come a point when the Lib Dems might say “No, enough”, and would pull out of the Coalition and force an election. I thought that maybe they’d be holding themselves back so that they could be really effective when something major cropped up. Something that, if not tackled, would be disastrous for Britain. Something that would be difficult to reverse once change had been made. Something like destruction of the NHS.
Well, it didn’t happen. I resigned from the party some months ago when it was clear how things were going, but I retained a vestige of hope that Lib Dems would, at the eleventh hour, find the changes too hard to swallow. I was encouraged by Evan Harris coming out strongly to say the things that needed saying. But no.
But if you didn’t resist on the basis of conscience and principle, I had thought you might have the sense to resist on more pragmatic grounds. Yesterday, a poll showed that 8% of the population would vote Lib Dem. I predict that will fall further after this weekend. Think ahead to the next election. If you were the party who stood up to the Conservatives and prevented them from wrecking the NHS, you’d gain a lot of kudos with your traditional supporters. But instead, you are the party who helped the Conservatives push through marketisation of the NHS. Well, there are many voters who may want marketisation, but they’re not going to vote for you at the next election, they’re going to vote Conservative. Your traditional voters didn’t want any of that, and will abandon you, as many, like me, already have. Baroness Williams argued that the Lib Dems have watered down the bill to make it more palatable. I’m sorry, people just aren’t going to vote for a party whose only role seems to be to help the Conservatives achieve their aims while not really believing in those aims. If you can’t see that, you’re not fit to be party leader.


Saturday, 10 March 2012

Blogging in the service of science

In my last blogpost, I made some critical comments about a paper that was published in 2003 in the Proceedings of the National Academy of Sciences (PNAS). There were a number of methodological failings that meant that the conclusions drawn by the authors were questionable. But that was not the only point at issue. In addition, I expressed concerns about the process whereby this paper had come to be published in a top journal, especially since it claimed to provide evidence of efficacy of an intervention that two of the authors had financial interests in.
It’s been gratifying to see how this post has sparked off discussion. To me it just emphasises the value of tweeting and blogging in academic life: you can have a real debate with others all over the world. Unlike the conventional method of publishing in journals, it’s immediate. But it’s better than face-to-face debate, because people can think about what they write, and everyone can have their say.
There are three rather different issues that people have picked up on.
1. The first one concerns methods in functional brain imaging; the debate is developing nicely on Daniel Bor’s blog and I’ll not focus on it here.
2. The second issue concerns the unusual routes by which people get published in PNAS. Fellows of the National Academy of Science are able to publish material in the journal with only “light touch” review.  In this article, Rand and Pfeiffer argue that this may be justified because papers that are published via this route include some with very high citation counts. My view is that the Temple et al article illustrates that this is a terrible argument. Temple et al have had 270 citations, so would be categorised by Rand and Pfeiffer as a “truly exceptional” paper. Yet, it contains basic methodological errors that compromise its conclusions. I know some people would use this as an argument against peer review, but I’d rather say this is an illustration of what happens if you ignore the need for rigorous review. Of course, peer review can go wrong, and often does. But in general, a journal’s reputation rests on it not publishing flawed work, and that’s why I think there’s still a role for journals in academic communications. I would urge the editors of PNAS, however, to rethink their publication policy sp that all papers, regardless of the authors, get properly reviewed by experts in the field. Meanwhile, people might like to add their own examples of highly cited yet flawed PNAS “contributions” to the comments on this blogpost.
3. The third issue is an interesting one raised by Neurocritic, who asked “How much of the neuroimaging literature should we discard?”  Jon Simons (@js_simons) then tweeted “It’s not about discarding, but learning”.  And, on further questioning, he added “No study is useless. Equally, no study means anything in isolation. Indep replication is key.”  and then “Isn't it the overinterpretation of the findings that's the problem rather than paper itself?” Now, I’m afraid this was a bit too much for me. My view of the Temple et al study was that it was not so much useless as positively misleading. It was making claims about treatment efficacy that were used to promote a particular commercial treatment in which the authors had a financial interest. Because it lacked a control group, it was not possible to conclude anything about the intervention effect. So to my mind the problem was “the paper itself”, in that the study was not properly designed. Yet it had been massively influential and almost no-one had commented on its limitations.
At this point, Ben Goldacre (@bengoldacre) got involved. His concerns were rather different to mine, namely “retraction / non-publication of bad papers would leave the data inaccessible.”  Now, this strikes me as a rather odd argument. Publishing a study is NOT the same as making the data available. Indeed, in many cases, as in this one, the one thing you don’t get in the publication is the data. For instance, there’s lots of stuff in Temple et al that was not reported. We’re told very little about the pattern of activations in the typical-reader group, for instance, and there’s a huge matrix of correlations that was computed with only a handful actually reported. So I think Ben’s argument about needing access to the data is beside the point. I love data as much as he does, and I’d agree with him that it would be great if people deposited data from their studies in some publicly available archive so nerdy people could pick over them. But the issue here is not about access to data. It’s about what do you do with a paper that's already published in a top journal and is actually polluting the scientific process because its misleading conclusions are getting propagated through the literature.
My own view is that it would be good for the field if this paper was removed from the journal, but I’m a realist and I know that won’t happen. Neurocritic has an excellent discussion of retraction and alternatives to retraction in a recent post,  which has stimulated some great comments. As he notes, retraction is really reserved for cases of fraud or factual error, not for poor methodology. But, depressing though this is, I’m encouraged by the way that social media is changing the game here. The Arsenic Life story was a great example of how misleading, high-profile work can get put in perspective by bloggers, even if peer reviewers haven’t done their job properly.  If that paper had been published five years ago, I am guessing it would have been taken far more seriously, because of the inevitable delays in challenging it through official publication routes. Bloggers allowed us to see not only what the flaws were, but also rapidly indicated a consensus of concern among experts in the field. The openness of the blogosphere means that opinions of one or two jealous or spiteful reviewers will not be allowed to hold back good work, but equally, cronyism just won’t be possible.  
We already have quite a few ace neuroscientist bloggers: I hope that more will be encouraged to enter the fray and help offer an alternative, informal commentary on influential papers as they appear.


Monday, 5 March 2012

Time for neuroimaging (and PNAS) to clean up its act

© www.CartoonStock.com

There are rumblings in the jungle of neuroscience. There’s been a recent spate of high-profile papers that have drawn attention to methodological shortcomings in neuroimaging studies (e.g., Ioannidis, 2011; Kriegeskorte et al., 2009; Nieuwenhuis et al, 2011) . This is in response to published papers that regularly flout methodological standards that have been established for years. I’ve recently been reviewing the literature on brain imaging in relation to intervention for language impairments and came across this example.
Temple et al (2003) published an fMRI study of 20 children with dyslexia who were scanned both before and after a computerised intervention (FastForword) designed to improve their language. The article in question was published in the Proceedings of the National Academy of Sciences, and at the time of writing has had 270 citations. I did a spot check of fifty of those citing articles to see if any had noted problems with the paper: only one of them did so. The others repeated the authors’ conclusions, namely:

1. The training improved oral language and reading performance.
2. After training, children with dyslexia showed increased activity in multiple brain areas.
3. Brain activation in left temporo-parietal cortex and left inferior frontal gyrus became more similar to that of normal-reading children.
4. There was a correlation between increased activation in left temporo-parietal cortex and improvement in oral language ability.
But are these conclusions valid? I'd argue not, because:
  • There was no dyslexic control group. See this blogpost for why this matters. The language test scores of the treated children improved from pre-test to post-test, but where properly controlled trials have been done, equivalent change has been found in untreated controls (Strong et al., 2011). Conclusion 1 is not valid.
  • The authors presented uncorrected whole brain activation data. This is not explicitly stated but can be deduced from the z-scores and p-values. Russell Poldrack, who happens to be one of the authors of this paper, has written eloquently on this subject: “…it is critical to employ accurate corrections for multiple tests, since a large number of voxels will generally be significant by chance if uncorrected statistics are used. .. The problem of multiple comparisons is well known but unfortunately many journals still allow publication of results based on uncorrected whole-brain statistics.” Conclusion 2 is based on uncorrected p-values and is not valid.
  • To demonstrate that changes in activation for dyslexics made them more like typical children, one would need to demonstrate an interaction between group (dyslexic vs typical) and testing time (pre-training vs post-training). Although a small group of typically-reading children was tested on two occasions, this analysis was not done. Conclusion 3 is based on images of group activations rather than statistical comparisons that take into account within-group variance. It not valid.
  • There was no a priori specification of which language measures were primary outcomes, and numerous correlations with brain activation were computed, with no correction for multiple comparisons. The one correlation that the authors focus on (Figure reproduced below) is (a) only significant on a one-tailed test at .05 level; (b) driven by two outliers (encircled), both of whom had a substantial reduction in left temporo-parietal activation associated with a lack of language improvement. Conclusion 4 is not valid. Incidentally, the mean activation change (Y-axis) in this scatterplot is also not significantly different from zero. I'm not sure what this means, as it’s hard to interpret the “effect size” scale, which is described as “the weighted sum of parameter estimates from the multiple regression for rhyme vs. match contrast pre- and post-training.”
Figure 2 from Temple et al. (2003). Data from dyslexic children.

How is it that this paper has been so influential? I suggest that it is largely because of the image below, summarising results from the study. This was reproduced in a review paper by the senior author that appeared in Science in 2009. This has already had 42 citations. The image is so compelling that it’s also been used in promotional material for a commercial training program other than the one that was used in the study. As McCabe and Castel (2008) have noted, a picture of a brain seems to make people suspend normal judgement.









I don’t like to single out a specific paper for criticism in this way, but feel impelled to do so because the methodological problems were so numerous and so basic. For what it’s worth, every paper I have looked at in this area has had at least some of the same failings. However, in the case of Temple et al (2003) the problem is compounded by the declared interests of two of the authors, Merzenich and Tallal, who co-founded the firm that markets the FastForword intervention. One would have expected a journal editor to subject a paper to particularly stringent scrutiny under these circumstances.
We can also ask why those who read and cite this paper haven’t noted the problems. One reason is that neuroimaging papers are complicated and the methods can be difficult to understand if you don’t work in the area.
Is there a solution? One suggestion is that reviewers and readers would benefit from a simple cribsheet listing the main things to look for in a methods section of a paper in this area. Is there an imaging expert out there who could write such a document, targeted at those like me, who work in this broad area, but aren’t imaging experts? Maybe it already exists, but I couldn’t find anything like that on the web.
Imaging studies are expensive and time-consuming to do, especially when they involve clinical child groups. I'm not one of those who thinks they aren't ever worth doing. If an intervention is effective, imaging may help throw light on its mechanism of action. However, I do not think it is worthwhile to do poorly-designed studies of small numbers of participants to test the mode of action of an intervention that has not been shown to be effective in properly-controlled trials. It would make more sense to spend the research funds on properly controlled trials that would allow us to evaluate which interventions actually work.

References
Gabrieli, J. D. (2009). Dyslexia: a new synergy between education and cognitive neuroscience. Science, 325(5938), 280-283.
Ioannidis, J. P. A. (2011). Excess significance bias in the literature on brain volume abnormalities. Arch Gen Psychiatry, 68(8), 773-780. doi: 10.1001/archgenpsychiatry.2011.28
Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F., & Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. [10.1038/nn.2303]. Nature Neuroscience, 12(5), 535-540. doi: http://www.nature.com/neuro/journal/v12/n5/suppinfo/nn.2303_S1.html 

McCabe, D., & Castel, A. (2008). Seeing is believing: The effect of brain images on judgments of scientific reasoning Cognition, 107 (1), 343-352 DOI: 10.1016/j.cognition.2007.07.017 

Nieuwenhuis, S., Forstmann, B. U., & Wagenmakers, E.-J. (2011). Erroneous analyses of interactions in neuroscience: a problem of significance. [10.1038/nn.2886]. Nature Neuroscience, 14(9), 1105-1107.

Poldrack, R. A., & Mumford, J. A. (2009). Independence in ROI analysis: where is the voodoo? Social Cognitive and Affective Neuroscience, 4(2), 208-213.

Strong, G. K., Torgerson, C. J., Torgerson, D., & Hulme, C. (2010). A systematic meta-analytic review of evidence for the effectiveness of the ‘Fast ForWord’ language intervention program. Journal of Child Psychology and Psychiatry, in press, doi: 10.1111/j.1469-7610.2010.02329.x.

Temple, E., Deutsch, G. K., Poldrack, R. A., Miller, S. L., Tallal, P., Merzenich, M. M., & Gabrieli, J. D. E. (2003). Neural deficits in children with dyslexia ameliorated by behavioral remediation: Evidence from functional MRI. Proceedings of the National Academy of Sciences of the United States of America, 100(5), 2860-2865. doi: 10.1073/pnas.0030098100



Friday, 24 February 2012

Neuroscientific interventions for dyslexia: red flags

I’m often asked for my views about interventions for dyslexia and related disorders. In recent years there has been a proliferation of interventions offered on the web, many of which claim to treat the brain basis of dyslexia. In theory, this seems a great idea; rather than slogging away at teaching children to read, fix the underlying brain problem. If your child is struggling at school, it can be very tempting to try something that claims to re-organise or stimulate the brain. The problem, though, is sorting the wheat from the chaff. There's no regulation of educational interventions and it can be hard for parents to judge whether it is worth investing time and money in a new approach.
My aim here is to provide some objective criteria that can be used. First, there is scientific evaluation: does the intervention have a plausible basis, and how has it been tested? Where claims are made about changing the brain, are they based on solid neuroscientific research? Second, there are red flags, some of which I listed in a previous post on ‘Pioneering treatment or quackery?” Here I've gathered these together so that there is a ready checklist that can be applied when a new intervention surfaces.

1. Who is behind the treatment and what are their credentials?
What you should look for here are relevant qualifications, particularly a higher degree (preferably doctorate) from an academic institution with a good reputation. Red flags are:
  • No information about who is involved ▶#1 
  • Intervention developed by someone with no academic credentials ▶#2 
  • Citation of spurious credentials; affiliation with organisations that have very lax membership criteria, e.g., Royal Society of Medicine ▶#3 
  • A lack of publications in peer-reviewed journals. Publications only in books counts as a red flag, because there is no quality control. ▶#4
It can be hard for a lay person to evaluate point #3, because some people cite qualifications that sound impressive but have no credibility. Academics in the field, however, can quickly identify whether a string of letters is indicative of prestige, or whether they are a smokescreen for lack of formal qualifications.
As far as #4 is concerned, relevant information can be obtained checking an author against a database such as Web of Science. However, access to such databases is largely restricted to academic institutions. Google Scholar is widely available, though its results are not restricted to peer-reviewed literature.

2. Is there a credible scientific basis to the treatment?
This is often difficult for a lay person to evaluate. Google Scholar may be helpful in tracking down articles that discuss the background to the intervention. Ideally, one is looking for a review by someone who is independent of those who developed it and who has good academic credentials. If no relevant journal articles are found on Google Scholar this is a red flag ▶#5. If a journal article is found, try to find whether the journal is a mainstream peer-reviewed publication.

3. Who is the intervention recommended for?

It is implausible that the whole gamut of neurodevelopmental problems has a single underlying cause, and it is unlikely that they will all respond to the same intervention. If an intervention claims to be effective for a host of diverse disorders, then this is a red flag ▶#6.

4. Is there evidence from controlled trials that the intervention is effective?
If there is such evidence, the main website for the intervention should describe it and provide links to the sources. No mention of controlled trials ▶#7, and heavy reliance on testimonials ▶#8 are both red flags. Chldren's progress should be measured on standardized and reliable psychometric tests, i.e. measures that have been developed for this purpose where normal range performance has been established. Failure to provide such information is another red flag ▶#9. It is not uncommon to find reference to trials with no controls, i.e. children’s progress is monitored before and after the intervention, and improvements are described. This is not adequate evidence of efficacy, for reasons I have covered in detail here: essentially, improvement in test scores can arise because of practice on the tests, maturation, statistical variation or expectation effects. If scores from before and after treatment are presented as evidence for efficacy, with no reference to control data, this is another red flag ▶#10, because it indicates that the practitioners do not understand the basic requirements of treatment evaluation.
If the evidence comes solely from children tested by people with a commercial interest, there may even be malpractice, with scores massaged to look better than they are. When there were complaints about an US intervention, Learning RX, ex-employees claimed that they had been encouraged to alter children's test scores to make their progress look better than it was (see comment from 6th Dec 2009). One hopes this is not common, but it is important to be alert to the possibility and to ensure those administering psychological tests are appropriately qualified, and if necessary get an independent assessment.
The strongest evidence for effectiveness comes from randomised controlled trials, which adopt stringent methods that have become the norm in clinical medicine. Where several trials have been conducted, then it is possible to combine the findings in a systematic review, which uses rigorous standards to evaluate evidence to avoid bias that can arise if there is ‘cherrypicking’ of studies. This level of evidence is very rare in behavioural interventions for neurodevelopmental disorders because the studies are expensive and time-consuming to do.

5. What is the attitude of those promoting the intervention to conventional approaches?
The question that an advocate for a new treatment has to answer is, if this is such a good thing, why hasn’t it been picked up by mainstream practitioners?
An answer that implies some kind of conspiracy by the mainstream to suppress a new development is a red flag ▶#11. This kind of argument is widely used by alternative medicine practitioners who maintain that others have vested interests (e.g. payments from pharmaceutical companies). This doesn’t hold water: basically, if a treatment is effective, then it makes no financial sense to reject it, given that people will pay good money for something that works.
Another red flag ▶#12 is if the new intervention is promoted alongside other alternative medicine methods that do not have good supportive evidence. This suggests that the practitioners do not take an evidence-based approach.


6. Are the costs transparent and reasonable?
Lack of information about costs on the website is a red flag ▶#13, especially if you can only get information by phoning (hence allowing the practitioner to adopt a hard sell approach). Is there any provision for a refund if the intervention is ineffective? If someone tells you their treatment has a 90% success rate, then they should be willing to give you your money back if it doesn't work. Another red flag is if you are asked to sign up in advance for a long-term treatment plan ▶#14. For example, in the case of the Dore programme, there were instances of families tied into credit agreements and forced to pay even if they don’t continue with the intervention.  

I’ll illustrate by applying the criteria to Sensory Activation Solutions. This is just one example of neuroscientific interventions on offer on the web. I've singled it out because I was recently asked my opinion after a new SAS Centre opened in Milton Keynes this month.
1. Who is behind the treatment and what are their credentials?
The SAS website states Sensory Activation Solutions (SAS) is the 'brainwave' of Steven Michaëlis and Kaśka Gozdek-Michaëlis and is the culmination of over 30 years of study and work relating to how we learn and how we can be more effective in life. I tried various approaches to search terms but was not able to find any publications by either person on Google Scholar. This is worrying: one would expect 30 years of study to yield some peer-reviewed papers. 
The biography of Steven Michaëlis does not mention any academic qualifications. He has a background in sound processing and computer technologies and has trained as a group counsellor. The website states that: Kaśka Gozdek-Michaëlis is an inter-faith, cross-cultural lecturer, writer, psychotherapist and life-coach with over 25 years experience. She gained a Master Degree in Oriental Studies at the prestigious University of Warsaw, Poland. She is the author of two books in Polish, 'Develop your genius mind' and 'Super-possibilities of your mind'.
Overall, the originators of the treatment are up-front about their background and do not hide behind spurious qualifications. However, neither of them appears to have any training in brain science or neurodevelopmental disorders, and their methods have not been subject to peer review. Two red flags:▶#2 ▶#4 

2. Is there a credible scientific basis to the treatment?
There were no entries in Google Scholar for "Sensory activation solutions", so I read the section on The science behind the SAS programs. This provided a quite complex story, about how playing sounds through headphones "activates the auditory processing centres in the brain... leading to less sensory overload, faster understanding, better verbal expression and improved reading and writing." It is a truism that playing sounds to people will activate auditory centres of the brain: that's what hearing is. The key question is whether the sounds used by SAS do anything special. There are numerous components to the SAS package, including use of vision, touch, taste, smell and proprioception "to reduce sensory overload." Sensory overload is a problem for some children, notably a subset of those with autistic spectrum disorder. But it's not generally viewed as problematic for children with dyslexia. It's also claimed that by presenting different auditory stimuli to the left and right ears, the SAS method can promote right-ear dominance and inter-hemispheric integration. In a video presentation, Michaëlis explains this aspect of the theory further, picking up on some old ideas about cerebral lateralisation, interhemispheric communication and rapid auditory processing. To those who don't know the literature, this will sound convincing, but his account is oversimplistic, and makes leaps from theory to intervention with no evidence. For example, with current methods of imaging it would be possible to test whether SAS stimuli alter children's cerebral lateralisation, but there's no indication of any studies investigating this. Overall, the account of the brain bases of dyslexia is out of line with contemporary neuroscientific research. One red flag: ▶#5

3. Who is the intervention recommended for?
SAS is described as appropriate for attention deficit disorder,  hyper-activity,  dyslexia, dyscalculia, hearing and speech disorders,  stammering,  autism,  Asperger's Syndrome,  Down Syndrome,  global developmental delay, Cerebral Palsy, eating disorders, sleeping disorders. In the video it is also recommended for acquired aphasia. One red flag: ▶#6

4. Is there evidence from controlled trials that the intervention is effective?
The "research" section of the website cites descriptive statistics only, largely based on parent satisfaction indices. There is no evidence that psychometrically sound measures were used to evaluate progress. 
There is a small scientific literature on Auditory Integration training (AIT), which has many features in common with aspects of the SAS package; most  studies focussed on autism, where there is little evidence of efficacy (Sinha et al, 2006). The American Speech-Language-Hearing Association concluded a review of AIT thus: Despite approximately one decade of practice in this country, this method has not met scientific standards for efficacy and safety that would justify its inclusion as a mainstream treatment for these disorders. Four red flags: ▶#7 ▶#8 ▶#9 ▶#10

 5. What is the attitude of those promoting the intervention to conventional approaches?
The 'resources' section of the website contained a wealth of information about other kinds of intervention, both mainstream and alternative. 

 6. Are the costs transparent and reasonable? 
The website was quite complicated to navigate, and I may have missed something, but I could not find any information about costs of treatment, only a phone number. It's not possible therefore to say if costs are reasonable. It seems unlikely that clients would be tied in to long-term contracts, as treatment duration seems quite short, lasting weeks rather than months. One red flag: ▶#13

Overall, you can see that SAS earns nine red flags on my evaluation scale.

I suspect no intervention is perfect, and if you have a child who is struggling at school you may want to go ahead and try an intervention regardless of red flags. My goal here is not to stop people trying new interventions, but to ensure that they do so with their eyes open. If practitioners make claims about changing the brain, then they can expect to have those claims scrutinised by neuroscientists. The list of red flags is intended to help people make informed decisions: it may also serve the purpose of indicating to practitioners what they need to do to win confidence of the scientific community.  

Reference
Sinha, Y., Silove, N., Wheeler, D., & Williams, K. (2006). Auditory integration training and other sound therapies for autism spectrum disorders: a systematic review Archives of Disease in Childhood, 91 (12), 1018-1022 DOI: 10.1136/adc.2006.094649

P.S. 6th March 2013: Here are some additional tips for spotting bad science more generally:

Sunday, 29 January 2012

2011 Orwellian Prize for Journalistic Misrepresentation

© www.CartoonStock.com
So the time has come round for the announcement of the 2011 Orwellian Prize. The prize is given for an article in an English-language national newspaper that achieves an unusually high level of inaccuracy. Only articles that describe a piece of published scientific research are eligible. Points are given for every statement in the article that does not match the original source, as follows:
  •     Factual error in the headline: 3 points
  •     Factual error in a subtitle: 2 points
  •     Factual error in the body of the article: 1 point
Last year, I got two nominations, but, as described here, neither adequately met the criteria. This year, I’ve had just one nomination, from Neurobonkers, but it’s set a standard for inaccuracy that will be hard to beat. The article that I first selected in 2010 to illustrate the scoring system scored 16 points. This one achieves a startling 23 points. The source article by Kucewicz et al (2011) can be found here. Here is a screenshot of the account in the Daily Mail, with errors marked in red (3 points), orange (2 points) and blue (1 point).
There is a detailed analysis of errors in this blogpost by Neurobonkers, which I urge you to look at. Suffice it to say,  the academic paper is not about cannabis, smoking or schizophrenia. Rather it is about an artificial compound that is not present in cannabis, which was injected into rats, and which led to changes in their brain waves.

There were some complaints to the Press Complaints Commission, and presumably in response to this, the article was modified. The headline, which originally read Just ONE cannabis joint 'can bring on schizophrenia' as well as damaging memory was altered to Just ONE cannabis joint 'can cause psychiatric episodes similar to schizophrenia' as well as damaging memory. Perhaps even the Daily Mail found the notion of a schizophrenic rat implausible. But the rest of the article remains, as a scare story about cannabis. And here is what makes this article such a prime candidate for the Orwellian Award: this is not about a hyped press release by a university, or misunderstanding of complex science. It's not even about sensationalising a scientific finding to draw readers in. No, this is about using a scientific paper as a prop in the Daily Mail's anti-cannabis campaign. A ploy that the newspaper has previously used in another ideological battle, on climate change. When reporting research, no respect is given to the truth: scientists are simply used to bolster a preconceived opinion, and if they don't do that, their findings are distorted.

Twelve of this article's 23 points came from the headline. Journalists don’t write the headlines. They therefore dislike my scoring system because it penalises errors in headlines more than errors in the body of the text. My view is that it’s the headlines that count for most: far more people will read the headline than the text, and for many readers it's the only part of the article they will process. It’s important that it's accurate. Although few would defend frank lying, some editors seem to think it doesn’t matter if a headline is hyperbolic, provided it sells the paper or gets someone to read further. This very issue has been a topic of debate in the recent Leveson Inquiry into culture, practice and ethics in the UK media. I feel strongly that it's a cop-out to just wash one's hands of it and blame anonymous sub-editors for misleading headlines. I shall therefore continue to award points in proportion to the prominence of the material. But I appreciate it’s not then fair to make the award to the journalist. Indeed, given the Mail's agenda on cannabis, the journalist in this case may well have been under duress to write a scare story.   I will accordingly be making the award to Paul Dacre, Editor of the Daily Mail. I will be happy also to send the token of appreciation for the nomination to Neurobonkers if he/she is willing to email me to tell me where to send it.

I'm pleased not to have had more nominations this year: it suggests that, despite all the grumblings about science journalism, the field is in rude health. I've certainly read a lot of good science reportage in our national newspapers, and where articles have made me angry, it's often because of hype by a press office or scientist, rather than distortion by the press. There are, however, still a few topics, among them drugs policy, where the political stakes are high and scientific reporting is cynically exploited to support an otherwise weak argument.

Coming up with an award certificate and item turns out to be an excellent way of avoiding doing serious work.....




Reference
Kucewicz, M., Tricklebank, M., Bogacz, R., & Jones, M. (2011). Dysfunctional Prefrontal Cortical Network Activity and Interactions following Cannabinoid Receptor Activation Journal of Neuroscience, 31 (43), 15560-15568 DOI: 10.1523/JNEUROSCI.2970-11.2011