Showing posts with label bias. Show all posts
Showing posts with label bias. Show all posts

Sunday, 23 April 2017

Sample selection in genetic studies: impact of restricted range


I'll shortly be posting a preprint about methodological quality of studies in the field of neurogenetics. It's something I've been working on with a group of colleagues for a while, and we are aiming to make recommendations to improve the field.

I won't go into details here, as you will be able to read the preprint fairly soon. Instead, what I want to do here is to expand on a small point that cropped up as I looked at this literature, and which I think is underappreciated.

It's to do with sampling. There's a particular problem that I started to think about a while back when I heard someone give a talk about a candidate gene study. I can't remember who it was or even what the candidate gene was, but basically they took a bunch of students, genotyped them, and then looked for associations between their genotypes and measures of memory. They were excited because they found some significant results. But I was, as usual, sitting there thinking convoluted thoughts about all of this, and wondering whether it really made sense. In particular, if you have a common genetic variant that has such a big effect on memory, would this really show up in a bunch of students – who are presumably people who have pretty good memories? Wouldn't it rather be the case that what you'd expect would be an alteration in the frequencies of genotypes in the student population?

Whenever I have an intuition like that, I find the best thing to do is to try a simulation. Sometimes the intuition is confirmed, and sometimes things turn out different and, very often, more complicated.

But this time, I'm pleased to say my intuition seems to have something going for it.

So here's the nuts and bolts.

I simulated genotypes and associated phenotypes by just using R's nice mvrnorm function. For the examples below, I specified that a and A are equally common (i.e. minor allele frequency is .5), so we have 25% as aa, 50% as aA, and 25% AA. The script lets you specify how closely these are related to the phenotype, but from what we know about genetics, it's very unlikely that a common variant would have a value more than about .25.

We can then test for two things:
1)  How far does the distribution of genotypes in the sample (i.e. people who are aa, aA or AA) resemble that in the general population? If we know that MAF is .5, we expect this distribution to be 1:2:1.
2) We can assign each person a score corresponding to number of A alleles (coding aa as zero, aA as 1, and AA as 2) and look at the regression of the phenotype on the genotype. That's the standard approach to looking for genotype-phenotype association.

If we work with the whole population of simulated data, these values will correspond to those that we specified in setting up the simulation, provided we have a reasonably large sample size.

But what if we take a selective sample of cases who fall above some cutoff on the phenotype? This is equivalent to taking, for instance, a sample from a student population from a selective institution, when the phenotype is a measure of cognitive function. You're not likely to get into the institution unless you have a good cognitive ability. Then, working with this selected subgroup, we recompute our two measures, i.e. the proportions of each genotype, and the correlation between the genotype and the phenotype.

Now, the really interesting thing here is that, as the selection cutoff gets more extreme, two things happen:
a) The proportions of people with different genotypes starts to depart from the values expected for the population in general. We can test to see when the departure becomes statistically significant with a chi square test.
b) The regression of the phenotype on the genotype weakens. We can quantify this effect by just computing the p-value associated with the correlation between genotype and phenotype.

Figure 1: Genotype-phenotype associations for samples selected on phenotype

Figure 1 shows the mean phenotype scores for each genotype for three samples: an unselected sample, a sample selected with z-score cutoff zero (corresponding to the top 50% of the population on the phenotype) and a sample selected with z-score cutoff of .5 (roughly selecting the top third of the population).

It's immediately apparent from the figure that the selection dramatically weakens the association between genotype and phenotype. In effect, we are distorting the relationship between genotype and phenotype by focusing just on a restricted range. 

Comparison of p-values from conventional regression approach and chi square test on genotype frequencies in relation to sample selection

Figure 2 shows the data from another perspective, by considering the statistical results from a conventional regression analysis, when different z-score cutoffs are used, selecting an increasingly extreme subset of the population. If we take a cutoff of zero – in effect selecting just the top half of the population, the regression effect (predicting phenotype from genotype), shown in the blue line, which was strong in the full population, is already much reduced. If you select only people with z-scores of .5 or above (equivalent to an IQ score of around 108), then the regression is no longer significant. But notice what happens to the black line. This shows the p-value from a chi square test which compares the distribution of genotypes in relation to expected population values in each subsample. If there is a true association between genotype and phenotype, then greater the selection on the phenotpe, the more the genotype distribution departs from expected values. The specific patterns observed will depend on the true association in the population and on the sample size, but this kind of cross-over is a typical result.

So what's the moral of this exercise? Well, if you are interested in a phenotype that has a particular distribution in the general population, you need to be careful when selecting a sample for a genetic association study. If you pick a sample that has a restricted range of phenotypes relative to the general population, then you make it less likely that you will detect a true genetic association in a conventional regression analysis. In fact, if you take a selected sample, there comes a point when the optimal way to demonstrate an association is by looking for a change in the frequency of different genotypes in the selected population vs the general population.

No doubt this effect is already well-known to geneticists, and it's all pretty obvious to anyone who is statistically savvy, but I was pleased to be able to quantify the effect via simulations. It is clear that it has implications for those who work predominantly with selected samples such as university students. For some phenotypes, use of a student sample may not be a problem, provided they are similar to the general population in the range of phenotype scores. But for cognitive phenotypes that's very unlikely, and attempting to show genetic effects in such samples seems a doomed enterprise.

The script for this simulation, simulating genopheno cutoffs.R should be available here: 
https://github.com/oscci/SQING_repo

(This link updated on 29/4/17).






Saturday, 18 February 2017

The alt-right guide to fielding conference questions


After watching this interview between BBC Newsnight's Evan Davies and Sebastian Gorka, Deputy Assistant to Donald Trump, I realised I'd been handling conference questions all wrong. Gorka, who is a former editor of Breitbart News, gives a virtuoso performance that illustrates every trick in the book for coming out on top in an interview: smear the questioner, distract from the question, deny the premises, and question the motives behind a difficult question. Do everything, in fact, except give a straight answer. Here's what a conference Q and A session might look like if we all mastered these useful techniques.

ED: Dr Gorka, you claim that you can improve children's reading development using a set of motor exercises. But the data you showed on slide 3 don't seem to show that.

SG: That question is typical of the kind of bias from people working at British Universities. You seem hell-bent on discrediting any view that doesn't agree with your own preconceived position.

ED: Er, no. I just wondered about slide 3. Is the difference between those two numbers statistically significant?

SG: Why are people like you so obsessed with trivial details? Here we are showing marvellous improvements in children's reading, and all you can do is to pick away at a minor point.

ED: Well, you could answer the question? Are those numbers significantly different?

SG: It's not as if you and your colleagues have any expertise in statistics. The last talk by your colleague Dr Smith was full of mistakes. She actually did a parametric test in a situation that called for a nonparametric test.

ED: But can we get back to the question of whether your intervention had a significant effect.

SG: Of course it did. It's an enormous effect. And that's only part of the data. I've got lots of other numbers that I haven't shown here. And if we got to slide 3, just look at those bars: the red one is much higher than the blue one.

ED: But where are the error bars?

SG: That's just typical of you. Always on the attack. Look at the language you are using. I show you all the results in a nice bar chart, and all you can do is talk about error. Don't you ever think of anything else?

ED: Well, I can see we aren't going to get anywhere with that question, so let me try another one. Your co-author, Dr Trump, said that the children in your study all had dyslexia, whereas in your talk you said they covered the whole range of reading ability. That's rather confusing. Can you tell us which version is correct?

SG: There you go again. Always trying to pick holes in everything we do. Seems you're just jealous because your own reading programs don't have anything like this effect.

ED: But don't you think it discredits your study if you can't give a straight answer to a simple question?

SG: So this is what we get, ladies and gentleman. All the time. Fake challenges and attempts to discredit us.

ED: Well, it's a straightforward question. Were they dyslexic or not?

SG: Some of them were, and some of them weren't.

ED: How many? Dr Trump said all of them were dyslexic.

SG: You'll have to ask him. I've got parents falling over themselves to get their children enrolled, and I really don't have time for this kind of biased questioning.

Chair: Thank you Dr Gorka. We have no more time for questions.

Monday, 9 April 2012

BBC's 'extensive coverage' of the NHS bill

Last month, there was a remarkable disconnect between what was being reported on BBC News outlets and what was concerning many members of the public on social media. The Health and Social Care Bill was passed by Parliament on 21st March, despite massive objections from many of those working in the NHS, and those members of the general public who were aware of the bill. Evidence of this concern was apparent from the fact that a petition with 486,000 signatures was presented to the Lords by Lord David Owen on 19th March, supporting his view that consideration of the Bill should be deferred until after the Risk Register had been published. There had also been a rally on 7th March attended by thousands of NHS workers. During the month of March, when there was still an opportunity of killing the bill if the Liberal Democrats had come out against it, there appeared to be very little coverage of it by the BBC. Only after the Bill had been passed, did the BBC seem willing to run it as a news item.
There has been a fair bit of commentary on the lack of coverage, with some suggesting there may have been a deliberate conspiracy to keep quiet because of political pressures and/or vested interests of BBC executives in private health providers.
http://salt-mine.net/blog/2012/03/23/complaint-bbcqt-no-question-nhs-bill/ 
http://storify.com/isobelweinberg/bbc-coverage-of-the-nhs-bill 
http://caiwingfield.com/cms/2012/03/police-suppression-of-peaceful-pro-nhs-protest-march-17th-2012/
http://socialinvestigations.blogspot.co.uk/2012/03/lord-patten-of-barnes-bridgepoint-and.html
Like many people, I submitted a complaint to the BBC about their lack of coverage of this important topic. I received a prompt reply as follows:
Dear Dorothy 
Thank you for contacting us regarding BBC News coverage. We understand you believe BBC News did not sufficiently report on the opposition to the Health and Social Care Bill. BBC News has reported extensively on the opposition to the Health and Social Care Bill across our news programmes and bulletins since the Bill was originally proposed. We have reported on the health, political and business dimensions of the debate during our flagship news programmes and news bulletins and have heard from politicians, NHS workers, public sector workers and members of the public alike, as well as from supporters of the bill. There have been numerous protests and demonstrations held in opposition to the Government’s proposals. Such shows of opposition have been varied in size and were spread across the different stages of the bill’s formation. We believe we have accurately and fairly reflected the nature of this opposition in our news coverage. While you were unhappy about the level of coverage given to this, the political opposition to the Bill culminated in the House of Commons emergency debate on 20 March. Accordingly, the Commons debate featured heavily in our news coverage on the day and was the lead story during our main news bulletins. The Health and Social Care bill has been one of the biggest UK stories over the past few months and we believe we have afforded it the appropriate level coverage in a fair and impartial manner, allowing viewers and listeners to make up their own minds on the matter at hand
The phrase ‘extensive coverage’ did not reflect my impressions. I am not glued to the BBC, but I am a regular listener to the Radio 4 Today Programme and I was not aware of the NHS bill receiving any coverage at all. I therefore submitted a follow-up complaint asking if they could please give me details of specific programmes when the BBC had covered the NHS bill during the month of March. Again, they replied promptly and courteously. And here is what they said (with relevant sections from the websites in blue):
Dear Dr Bishop
We understand that you would like details of when opposition to the Health and Social Care Bill was covered by the BBC. Opposition to the bill has been covered on various programmes across the BBC, for example; Newsnight, The Daily Politics, Today and BBC News Online. Opposition was also covered on 'Newsnight' during a report on 9th March which looked at Liberal Democrat activist's plans to derail the bill and on the 13th March during a discussion on the future of the welfare system. A report on the 'The Daily Politics' broadcast on 13th March (at 14:09) highlighted opposition from Labour as well as the Royal College of GPs. Diane Abbott said health professionals were still opposed to the Health and Social Care Bill, which could be days away from becoming law. She said a future Labour government would overturn the act and "unpick the worst of the damage". Liberal Democrat spokesman Lord Clement-Jones said the bill was "going to get more acceptance"
More coverage was given on the 'Today' programme on the following dates:  
Sat 10th March, 0712-0714 Liberal Democrat activists will decide this morning whether Nick Clegg will face a vote on the health bill. The BBC's Robin Bryant explains why that could be bad news for the government. 
Weds 21st March, 0840-0853 The government's controversial plans to change the NHS have passed their final hurdle in Parliament after 14 months of opposition and changes in both houses. Professor Chris Ham, chief executive of the health think-tank The King's Fund, looks at what we left with now.
In addition examples of coverage on BBC News Online include: 


So, if I have understood this right,during March, the Today Programme covered the story once, in an early two-minute slot, before the Bill was passed. Other items that morning included 4 minutes on a French theme park based on Napoleon, 6 minutes on international bagpipe day and 8 minutes on Jubilee celebrations.
 



free counters