Monday, 30 May 2011

Are our ‘gold standard’ autism diagnostic instruments fit for purpose?

In 1985, Simon Baron-Cohen, Alan Leslie and Uta Frith published a landmark paper entitled “Does the autistic child have a theory of mind?” It described a small study that Simon Baron-Cohen completed for his doctoral thesis which, according to Google Scholar, has been cited 2800 times. If the same paper were submitted for publication today, most journals would reject it. Why? The paper stated: “The 20 autistic children had been diagnosed according to established criteria (Rutter, 1978)”. Nowadays this would be deemed inadequate. Many editors and reviewers insist that studies of autism use two diagnostic procedures, the Autism Diagnostic Interview - Revised (ADI-R) and the Autism Diagnostic Observation Schedule - Generic (ADOS-G). According to the NIH National Database for Autism Research, requiring people to use these “gold standard” assessments will “help accelerate scientific discovery”. But use of these instruments adds hugely to the time and money costs of research. Small-scale studies of autism by PhD students have become nonviable, and large-scale genetic and epidemiological studies are bogged down by the need to spend hours just establishing the phenotype for each case. Researchers from countries where ADI-R and ADOS-G are not available are at a serious disadvantage. And, as I shall argue, the end result is not a clearcut diagnosis.

Autism has three key defining features: impairments in communication, social interaction and behavioural repertoire. The latter encompasses both repetitive behaviours such as stereotyped movements, and restricted interests, e.g., an obsessive fascination with aeroplanes. After autism was first described by Leo Kanner in 1943, the diagnosis quickly became popular, but there were concerns that it was over-used. There was a clear need to translate Kanner’s clinical descriptions into more objective diagnostic criteria. The first step was to develop checklists of symptoms, and these were included for the first time in the 1980 version of the Diagnostic and Statistical Manual of the American Psychiatric Association, DSM-III. However, this still left room for uncertainty: clinicians might, for instance, disagree about interpretation of terms such as: “Pervasive lack of responsiveness to other people”.

The Autism Diagnostic Interview (ADI) was designed to address this problem. It was first published in 1989, with a revised form, ADI-R, appearing in 1994. The ADI-R typically takes 1.5 to 3 hours to administer, and covers all the symptoms of autism and related conditions. Items are coded on a 4-point scale, from 0 (absent) to 3 (present in extreme form). The scores from a subset of items are then combined to give a total for each of the three autism domains, and a diagnosis of autism is given if scores on all three domains are above cutoffs, and onset was evident by 36 months. Validation of the ADI-R was carried out by comparing scores for 25 preschool children who had clinically diagnosed autism and 25 nonautistic children with intellectual retardation or language impairment.

Right from the outset, however, there was concern that diagnosis of autism should not be made on the basis of parental report alone. Some parents are poor informants. On the one hand, they may fail to remember key features of their child’s behaviour; on the other hand, their memories may have been coloured by reading about autism. Parental report therefore needs to be backed up by observation of the child. The Autism Diagnostic Observation Schedule (ADOS), published in 1989, was designed for this purpose. It exposes the child to a range of situations designed to elicit autistic features, and particular behaviours, such as eye contact, are then coded by a trained examiner. ADOS-G, a generic version, was published in 2000, and covers a wide age range, from toddlers through to adults.

ADI-R and ADOS-G quickly became the instruments of choice for autism diagnosis. It was generally appreciated that if we use standard instruments, researchers and clinicians should be able to communicate about autism with a fair degree of confidence that they are referring to individuals who meet the same diagnostic criteria.

There is, however, a downside. ADI-R and ADOS-G were designed to be comprehensive, but they were not designed to be efficient. As noted above, ADI-R takes up to 3 hours to administer and score. ADOS-G takes about 45 minutes. In addition, testers must be trained to use each instrument, and it may take some months to find a place on a training course. Each course lasts around one week, and the trainee then has to do further assessments which are recorded and sent for validation by experts. This process can easily add another 6 months. For anyone under time pressure, such as a doctoral student or grant-holder with research assistants on fixed term contracts, the training requirements can make a study impossible to do. Inclusion of both ADI-R and ADOS-G in a test protocol can double or treble the duration of a project, especially where it is necessary to travel to interview parents who may be available only during anti-social hours.

So, does autism diagnosis need to involve such a lengthy process? This was a question I was prompted to consider when I was asked to speak at a roundtable debate on diagnostic tools for autism at the International Meeting For Autism Research (IMFAR) in London in 2008. I concluded that the answer is almost certainly no. As Matson et al (2007) put it: “Some measures emphasize the fact that they are very detailed. We would argue that detail equals time. From a pragmatic perspective, our view is that a major priority should be to develop the balance between obtaining relevant information to make a diagnosis, while parsing out items that do not enhance that goal”. (p. 49). I was surprised when I first undertook ADI-R training to find that the interview included many items that did not feature in the final algorithm. When I (and others) queried this, we were told that the interview worked in its entirety, and to pull out selected items would disrupt a natural flow. Also, the non-algorithm items might be useful for diagnosing conditions other than autism. While both points may be important in a clinical setting, they have much less force for the poor graduate student who is doing a doctorate on, say, perceptual processing in autism, and only wants to use the ADI-R to confirm a diagnosis that has already been made by a clinician. There was a short form, I was told, but it was not recommended and should only be used by clinicians, not researchers. I was even more surprised to find how the algorithm was devised. I was familiar with discriminant function analysis, whereby you take a set of scores on two (or more) groups, and find the best weighted sum of scores to discriminate the groups. You can then use the correlations between items to drop items from the algorithm successively until you get to the point where accuracy of diagnostic assignment declines if further items are dropped. I had assumed that this kind of statistical data reduction had been used to identify the optimal set of items for identifying autism. I was wrong. It seemed that the items were selected on the basis of their match to clinical descriptions of autism, and no attempt had been made to test the efficiency of the algorithm by dropping items.

There is good reason to believe that a much shorter and simpler procedure would be feasible. In 1999, a study was published comparing diagnostic accuracy of the ADI with a 40-item screening questionnaire. The accuracy of the questionnaire was as good as the interview. My impression is that this finding did not lead to rejoicing at the prospect of a shorter diagnostic procedure, but rather alarm that diagnosis could be reduced to a trivial box-checking exercise. While I have some sympathy with that view, I feel these results should have made the researchers pause to consider whether a much shorter and more efficient approach to diagnosis might be feasible.

But the problems get worse. As Jon Brock recently pointed out on his blog, Kanner’s view of autism as a distinct syndrome is no longer accepted. It’s clear that autism symptoms can occur in milder form, and that they do not necessarily all go together. The broader term ‘Autism Spectrum Disorder’ (ASD) is nowadays used to encompass these cases. The schematic illustration in Figure 1 illustrates the diagnostic problem: one has to decide where to place boundaries on the figure to distinguish ASD from normality, when in reality, all three domains of the autism triad of symptoms shade into normality, with no sharp cutoffs. The ADI-R algorithm specifies whether or not you have autism, and does not give cutoffs for milder forms of ASD. The ADOS-G does have cutoffs for milder forms, but is inappropriate for detailed assessment of repetitive behaviours/restricted interests, so it is not watertight. This means that, when it comes to diagnosis of ASD, a great deal is left to clinician’s judgements. These, rather, than algorithm scores, are used to arrive at diagnoses.


Figure 1: Schematic of autism as a spectrum disorder: Circle correspond to areas of deficit, red = social impairment, blue = communication difficulties; yellow = repetitive behaviour/restricted interests, with depth of colour indicating severity of impairment. Individuals with all three features (centre of the figure) meet full diagnostic critieria for autism, but those falling outside this region, who have milder or partial difficulties are candidates for a diagnosis of autism spectrum disorder. Note, however, there is no clearcut boundary between autistic spectrum disorder and normal variation.
There were quite a few researchers at the IMFAR meeting who complained that their papers were rejected by journals because they relied on a diagnosis made by an expert, rather than ADI-R and ADOS-G. They would be baffled to see that several recent state-of-the-art epidemiological studies use consensus judgement by expert clinicians to make their diagnoses - and these diagnoses don’t necessarily agree with ADI-R and ADOS-G. So, for instance, Baird et al had 81 cases who met a consensus clinical diagnosis of childhood autism, but only 53 (65%) were classified as autistic by the algorithms of both the ADOS-G and the ADI-R. They also identified 77 children with consensus diagnosis of ‘other ASD’ of whom 69% met criteria for autism on the ADI-R, and 38% met cutoff for PDD or autism on the ADOS. On the ADOS-G, 10% of non-autistic children scored above cutoff for either ASD or autism. This fits my experience: high scores can reflect lack of engagement, shyness, or language difficulties. Similarly, Baron-Cohen et al identified four cases of autism and seven with other ASDs in a population screening of children aged 5 to 9 years in Cambridgeshire. All the autism cases met autism criteria on both ADI-R and ADOS-G. Of the other ASD cases, five met criteria for autism on ADI-R but not ADOS-G, and two did not meet criteria for autism or ASD on either instrument. In describing these findings, I am not criticising the authors of these studies, whose methods were transparently reported and are consistent with practice as it has evolved in the field. But it is ironic that we seem to have come full circle. ADI-R and ADOS-G were developed to make diagnosis more objective, but because they aren’t geared up to diagnose ASD, we are thrown back on ‘expert clinical opinion’. This is far from reassuring, given that a recent study reported that, after months of training, researchers agreed well on scoring standardized instruments, but “consistent differences between sites in overall clinical impression were reported”.

My own recommendation is for a two-step procedure. The first step would involve a much briefer version of the ADI-R, which would be designed to pick up clear-cut cases of autism that everyone would agree on and distinguish them from clearly non-autistic cases. It’s an empirical question, but I suspect that if we were to do a stepwise discriminant analysis to identify a minimum set of diagnostic items, this would be considerably shorter than the current set used in ADI-R. The interview may then need redesigning so that it still flows fluently and follows a logical course, but this should not be impossible. In clinical settings, those identified might require further direct assessment to confirm diagnosis and identify specific needs, but this would not be necessary for determining who should be included in a research study. This would leave a group of children in whom autism was suspected but not confirmed. The question here is whether we will ever arrive at a diagnostic procedure that will clearly separate such children into ASD and non-ASD. I was involved in a study with just such a group of ‘marginal’ cases a few years ago, where we administered ADI-R and ADOS-G. The results were all over the place: some children looked autistic on ADI-R and not on ADOS-G and others showed the opposite pattern. Some had evidence of marked change in behaviour between preschool and school-age years. When I asked an autism expert how such cases should be categorised, he suggested I get an expert clinical opinion. Yet expert clinical opinion is not seen as adequate by many journal editors! And there is documented evidence that even experienced clinicians will disagree in cases where the child has a confusing pattern of symptoms, and that expert diagnoses are not stable over time. My suggestion is that in our current state of knowledge it makes no sense to try and get reliable cutoffs for identifying ASD. Instead we should aim to assess the nature and degree of impairments in different domains. Assessments such as the 3Di or Social Responsiveness Scale, which treat autistic features as dimensions rather than all-or-none symptoms, seem better suited to this task than the existing gold standards.

Finally, I must emphasise that, although I think the ADI-R and ADOS-G are not optimal for diagnosing ASD for research purposes, they nevertheless have value. They distill a great deal of clinical wisdom in the assessment process, and are cleverly crafted to pinpoint the key features of autism. Anyone who undergoes training in their use will come away with a far greater understanding of autism than they had when they started. However, these instruments are not well suited for addressing the NIH aim of “accelerating scientific discovery”. In research contexts they have the opposite effect, by making researchers go through an unnecessarily long and complex diagnostic process which does not yield suitable quantitative results for assessing the dimensional aspect of ASD.

Rondeau E, Klein LS, Masse A, Bodeau N, Cohen D, & Guilé JM (2010). Is Pervasive Developmental Disorder Not Otherwise Specified Less Stable Than Autistic Disorder? A Meta-Analysis. Journal of autism and developmental disorders PMID: 21153874

Wednesday, 25 May 2011

Scientific communication: the Comment option


















1. Think of something interesting to say

2. Hit the link saying ‘comment’.

3. Find you have to register. Enter first name, surname, title, job description, qualifiations, place of work, phone number, fax number, address.

4. Spend several minutes looking for England, then Great Britain on the drop-down list, before hitting on United Kingdom. Press enter, while wondering how often someone from Azerbaijan makes a comment.

4. Back to step 3 because failed to enter a zip code in the correct format.

5. Realise that the special security software that was installed to prevent unauthorised access to the computer has barred interaction with the site, and everything you have entered has been lost.

6. Select ‘allow this site’ in the security software.

7. Back to step 3.

8. Select a password.

9. Retype the password.

10. Retyped password fails to match password. Back to step 8.

11. Hooray! All entered.

12. Try to log in. System tells you someone with this email address has already registered with a different password.

13. Alarm husband with sudden outburst of profanities and table thumping.

14. Try to guess password.

15. Fail.

16. Loop back to step 14 several times.

17. Request new password.

18. Open email to look for message re new password.

19. No sign of email re password. Get distracted by other emails. Ponder on why as a neuropsychologist have been invited to contribute an article to Journal of Plant Membranes.

20. Delete messages with header ‘Dear Friend in Christ’. Don’t they realise I’m an atheist?

21. Password message appears in Inbox.

22. Back to login screen.

23. Login, only to find there is a saved password already.

24. Use of saved password gives error message.

25. Relogin with new password. Yes!

26. Search for original site for comments. Can’t find it.

27. Back to twitter to find the address of the original article that I wanted to comment on.

28. Find the article.

29. Hit the comment button.

30. Now what was it I wanted to say?

Monday, 16 May 2011

Autism diagnosis in cultural context

A review of Isabel’s World by Roy Richard Grinker



In 1993, when Roy Grinker’s daughter, Isabel, was two years old she did not talk or gesture, flapped her hands and arms, and did not make eye contact. At 32 months, she had mastered about 70 words, which she spoke clearly, but all were nouns and the list did not include ‘mommy’ or ‘daddy’. She didn’t say ‘yes’ or ‘no’ and pulled someone to the refrigerator when she was hungry. Nowadays it’s hard to imagine a paediatrician would fail to recognise classic symptoms of autism, but for the Grinkers the process of getting a diagnosis was protracted and painful. At best they met with ignorance, and at worst with professionals who were imbued with the ideas of Bruno Bettelheim and attributed Isabel’s problems to the fact that her mother went out to work. Once a diagnosis was obtained, the Grinkers still had to fight for appropriate educational provision. They came up against people who had no experience of autism and were unwilling or unable to engage with Isabel. Gradually, though, the battle was won. Isabel, whose high nonverbal ability was eventually recognised, was able to attend a mainstream school with support. Now in her teens, she still has major problems with communication and social interaction, and her parents accept she will not able to live independently. She has, however, found a niche in her community where she has friends, can enjoy her interests in music and animals, and is accepted for who she is.

Grinker’s book is more than a parent’s first-hand account of his daughter’s autism: it is also influenced by the fact that he is a cultural anthropologist, with a special interest in attitudes to health and disability in different parts of the world. He uses his professional insights to comment on two related issues: the rise in autism diagnosis in the USA, and attitudes to autism in other cultures, particularly France, Korea, South Africa and India.

Figure 1
The rise in autism diagnosis is a hotly debated topic. In general, there is agreement about the facts: the number of children receiving an autism diagnosis in the USA has risen sharply since the 1970s. The data in Figure 1 were taken from a website relating to the Individuals with Disability Education Act (IDEA), and they show a more than four-fold increase in children identified with autism in the ten year period between 1997 and 2006. I have previously noted in a PLOS One article that research on autism, as assessed by funding and publications, has skyrocketed over the same period.


The debate is over the reason for the increase: is it due to some environmental factor that causes autism, or can it be explained simply in terms of changes in diagnostic criteria and related factors? Grinker comes down solidly in favour of the latter explanation, and a substantial part of the book is devoted to tracing the history of autism diagnosis and related socio-cultural factors. A major turning point was the publication in 1987 of the third revision of the diagnostic and statistical manual of the American Psychiatric Association, DSM-III-R. Grinker cites a study of 194 children suspected of autism which found that whereas 51% met criteria on DSM-III, 91% met criteria on DSM-III-R. He also reports a typographic error that was made in the publication of DSM-IV in 1994 that led to the ‘autism spectrum’ disorder of PDDNOS being defined on the basis of a child having impairments in social interaction or verbal/nonverbal communication skills, when it should have required impairment in both domains.

Grinker tackles a question often asked by those who think the autism epidemic is genuine: if it’s just due to a change in diagnostic practices, where were all the undiagnosed autistic children in the past? Surely we would have noticed them? On the basis of a follow-up of a small sample of UK children with severe language impairments, our group found cases who would now be regarded as unambiguously autistic, but who were diagnosed as language-impaired when they were seen in the 1980s, prior to DSM-IIIR. So are we just engaging in 'diagnostic substitution', whereby the same child who is now regarded as autistic was previously given a different label? If so, we'd expect to see that as autism diagnosis goes up, language disorder diagnosis goes down. This would be consistent with an analysis of doctor's records in the UK, which found that the proportion of children with diagnoses such as speech/language disorder went down over the same period as autism diagnoses went up. In the IDEA data, however, I found no evidence for such a process: over the 1997-2006 time period, there's actually a slight increase in diagnoses of speech/language diagnoses. Grinker, however, suggests that, in the US, children with autism would, in the past, often have been diagnosed with mental retardation. Consistent with this, the IDEA data do show a corresponding decrease in diagnosis of mental retardation over the same time period as autism diagnoses increase (see Figure 1).

Another point stressed by Grinker is that a diagnosis of autism has consequences for the child’s access to services. A diagnosis may get your child Medicaid waivers so that they can receive a host of interventions at reduced cost. Grinker describes how he lost hundreds of dollars because a speech pathologist who worked with Isabel submitted the bills under the diagnosis of “Mixed Receptive-Expressive Language Disorder”. When the diagnosis was changed to autism, the bills were automatically reimbursed. There is therefore considerable pressure on paediatricians to give a diagnosis of autism rather than some other condition.

The implications of an autism diagnosis for intervention is very different in some other countries. In France, where the legacy of psychoanalysis has been long-lasting, autism is seen as a psychodynamic disorder and there is little educational provision for affected children. In India, the diagnosis is seldom made, even if doctors recognise autism, because there seems no point: there are no facilities for affected children. Grinker has particular interest in South Korea, his wife’s birthplace, where the stigma attached to disability is so great that children with autism will be hidden away, because otherwise their siblings’ marriage prospects will be blighted. Education is seen as the route to success in Korea, and it is normal for children to spend hours after school being coached; a child who struggles at school or who does not conform to expected standards of behaviour brings shame to the family.

Grinker’s book shows that the impact of labels is immense, even if they are worryingly arbitrary. It seems crazy, for instance, that a child with educational difficulties will get insurance cover for interventions, or special help at school on the basis of a diagnosis of autism, whereas a child with equally serious needs who is diagnosed with mental retardation or receptive language disorder gets nothing. But on the positive side, he notes how the growing awareness and acceptance of autism has brought huge benefits to children and their families. Just as in Korea, it used to be common for people in the US to be baffled or even frightened by autism, and for schools to shun children with any kind of disability. The landscape has changed massively over the past twenty years. Grinker notes how Isabel’s schoolmates look out for her, and how her presence in a mainstream classroom has beneficial effects on all pupils. Cultural stereotypes are being challenged, so much so that it is no longer regarded as amazing if a university student tells you they have autism.

This brings Grinker to a final point: which kinds of cultural setting are most supportive for people with autism? He presents some fascinating data showing that rural communities are usually far more positive than urban ones for people with all kinds of disability. In a small community, everyone will know the person with a disability, and see them as an individual. In the context of neurodevelopmental disorders in general, I’ve long been advocating that we need an educational system that doesn’t just try to ‘fix’ children, but rather identifies activities they enjoy and can do well, as a basis for finding them a niche in society.  This view is cogently expressed by Grincher “… in order to help people with autism we don’t always need to fully mainstream them, or pretend that they are not different, and we don’t need to simply reduce stigma. Rather, we need to provide roles in our communities for pepole with autism, some of which they may, in fact, be able to perform better than anyone else…” (p. 342).


Bishop, D., Whitehouse, A., Watt, H., & Line, E. (2008). Autism and diagnostic substitution: evidence from a study of adults with a history of developmental language disorder Developmental Medicine & Child Neurology, 50 (5), 341-345 DOI: 10.1111/j.1469-8749.2008.02057.x

ResearchBlogging.org

Wednesday, 11 May 2011

The X and Y of sex differences

Why and how are men and women different? My interest in this topic is fuelled by my research on neurodevelopmental disorders of language and literacy that typically are much more common in males than females. In this post, I am ranging far from my comfort zone in psychology to discuss what we know from a genetic perspective. My inspiration was a review in Trends in Genetics by Wijchers and Festenstein called “Epigenetic regulation of autosomal gene expression by sex chromosomes”. Despite the authors' sterling efforts to explain the topic clearly, I suspect their paper will be incomprehensible to those without a background in genetics, so I'll summarise the main points - with apologies to the authors if I over-simplify or mislead.

So, to start with, some basic facts about chromosomes in humans:
•    We have 23 pairs of chromosomes, one member of each pair inherited from the father, and the other from the mother.
•    For chromosome pairs 1-22, the autosomes, there is no difference between males and females.
•    Chromosome pair 23 is radically different for males and females: females have two X chromosomes, whereas males have an X chromosome paired with a much smaller Y chromosome
•    The Y chromosome carries a male-determining gene, SRY, which causes testes to develop. The testes produce male hormones which influence body development to produce a male.
•    The X chromosome contains over 1000 genes, compared to 78 genes on the Y chromosome.
•    In females, only one X chromosome is active. The other is inactivated early in development by a process called methylation. This leads to the DNA being formed into a tight package (heterochromatin), so genes from this chromosome do not get expressed. X-inactivation randomly affects one member of the X-chromosome pair early in embryonic development, and all cells formed by division of an original cell will have the same activation status. The patches of orange and black fur on a calico cat arise when a female has different versions of a gene for coat colour on the two X chromosomes, so patches of orange and black fur occur at random.
•    In both X and Y chromosomes, there is a region at the tip of the chromosome called the pseudoautosomal region, which behaves like an autosome, i.e., it contains homologous genes on X and Y chromosomes, which are not inactivated, and which recombine during the formation of sperm and eggs.
•    In addition, a proportion of genes on the X chromosome (estimated around 20%) escape X-inactivation, despite being outside the pseudoautosomal region.

These basic facts are summarised in Figure 1. Genes are symbolised by red dots, grey shading denotes an inactivated region, and yellow is the pseudoautosomal region.

Figure 1
Note that because (a) the male Y chromosome has few genes on it and (b) one X chromosome is largely inactivated in females, normal males (XY) and females (XX) are quite similar in terms of sex chromosome function: i.e., most of the genes that are expressed will come from a single active X chromosome.
Studies of mice and other species have, however, demonstrated differences in gene expression between males and females, and these affect tissues other than the sex organs, including the brain. Most of these sex differences are small, and it is usually assumed that they are the result of hormonal influences. Thus the causal chain would be that SRY causes the testes to develop, the testes generate male hormones, and those hormones affect how genes are expressed throughout the body.

You can do all kinds of things to mice that you wouldn't want to do to humans. For a start you can castrate them. You can then dissociate the effect of the XY genotype from the effect of circulating hormones. When this is done, many of the sex differences in gene expression disappear, confirming the importance of hormones.
There's some evidence, though, that this isn't the whole story. For a start, it is possible to find genes that are differently expressed in males and females very early in development, before the sex organs are formed. These differences can't be due to circulating hormones. You can go further and create genetically modified mice in which chromosome status and biological sex are dissociated.  For instance, if Sry (the mouse version of SRY) is deleted from the Y chromosome, you end up with a biologically female mouse with XY chromosome constitution. Or an autosomal Sry transgene can be added to a female to give a male mouse with XX constitution. A recent study using this approach showed that there are hundreds of mouse genes that are differently expressed in normal XX females vs. XY females, or in normal XY males vs. XX males. For these genes, there seems to be a direct effect of the  X or Y chromosome on gene expression, which isn't due to hormonal differences in males and females.

Wijchers and Festenstein consider four possible mechanisms for such effects.
1. SRY has long been known to be important for development of testes, but that does not rule out a direct role of this gene in influencing development of other organs. An in mice there is indeed some evidence for a direct effect of Sry on neuronal development.
2. Imprinting of genes on the X chromosome. This is where it starts to get really complicated. We have already noted how genes on the X chromosome can be inactivated. I've told you that X-inactivation occurs at random, as illustrated by the calico cat. However, there's a mechanism known as imprinting whereby expression of a gene depends on whether the gene is inherited from the father or the mother. Imprinting was originally described for genes on the autosomes, but there's considerable interest in the idea of imprinting affecting genes on the X chromosome, as this could lead to sex differences. The easiest way to explain this is by a mouse experiment. It's possible to make a genetically modified mouse with a single X chromosome. The interest is then in whether the single X chromosome comes from the mother or the father. And indeed, there's growing evidence for differences in brain development and cognitive function between  genetically modified mice with a single maternal or paternal X chromosome: i.e., evidence of imprinting. Now this has implications for sex differences in normal, unmodified mice. XY male mice have just one X chromosome, which always comes from the mother, and will always be expressed. But XX female mice have a mixture of active maternal and paternal X-chromosomes. Any effect that is specific to a paternally-derived X-chromosome will therefore only be seen in females.
What about humans? Here we can study girls with Turner syndrome, a condition in which there is one rather than two X chromosomes. Skuse and colleagues found differences in cognition, especially social functioning between girls with a single maternal X vs. those with a single paternal X. There are few studies of this kind, because it is difficult to recruit large enough samples, so the results need replicating. But potentially this finding has tremendous implications, not just for finding out about Turner syndrome itself, but for understanding sex differences in development and disorders of social cognition.
3. Although most X-chromosome genes are expressed from only one X-chromosome, as noted above, some genes escape inactivation, and for these genes females have two active copies. In the main, these are genes with a homologue on the Y-chromosome, but there are exceptions, and in such cases females have twice the dosage of gene product compared to males (see Figure 1). And even where there is a homologue gene on the Y chromosome, this may have different effects from the active X-chromosome gene.
4. The Y chromosome contains a lot of inactive DNA with no genes. Recent studies on fruit flies has found that this inactive DNA can affect expression of genes on the autosomes, by affecting availability in the cell nucleus of factors that are important for gene expression or repression. It's not clear if this applies to humans.

My interest in this topic has led me to study children who do not inherit the normal complement of sex chromosomes. These include girls with a single X chromosome (XO, Turner syndrome), girls with three X chromosomes (triple X or XXX syndrome), (see figure 2) and boys with an extra X (XXY or Klinefelter’s syndrome), and boys with an extra Y (XYY syndrome).(see Figure 3).
Figure 2

Figure 3
Affected children typically don’t have intellectual disability and attend regular mainstream school. As illustrated in Figures 2 and 3 , this makes sense because the genetic differences between those with missing or extra sex chromosomes and those with the normal XX or XY complement are not great. In Turner syndrome there is only one X chromosome, whereas children with XXX or XXY will have all but one X chromosome inactivated. The extra Y in boys with XYY contains only a few genes.

Nevertheless, although children with atypical sex chromosomes are not severely handicapped, distinctive neuropsychological profiles have been described. Girls with Turner syndrome often have poor visuospatial function and arithmetical ability, whereas language skills are typically impaired in children with an extra sex chromosome. To explain these effects, researchers have proposed a role for genes that normally escape inactivation, which will be underexpressed in Turner syndrome, or over-expressed in children with three sex chromosomes (sex chromosome trisomy) - see point 3 above.

Wijchers and Festenstein note the importance of individuals with sex chromosome anomalies for informing our understanding of sex chromosome effects on development, but their account is not very satisfactory, as they state that “females with triple X syndrome (47,XXX) seem normal in most cases.” Although it is the case that many girls with XXX go undetected, surveys of prenatally or neonatally identified cases indicate that they have cognitive problems. Language deficits are found at high levels in all three cases of trisomy, XXX, XXY and XYY, with a trend for lower overall IQ in girls with XXX than boys with XXY or XYY. We did a study based on parental report, and found a diagnosis of autism spectrum disorder was more common in boys with XXY and XYY than in boys with normal XY chromosome status. But there was considerable variability from child to child, with some having no evidence of any educational or social difficulties, and others with more serious learning difficulties or autism. We currently lack data that would allow us to relate the cognitive profile in such children to their detailed genetic makeup, but this in an area researchers are starting to explore. We are optimistic that such research will not only be helpful in predicting which children are likely to need additional help, but also may throw light on more global questions about the genetic basis of sex differences in cognitive abilities and disabilities.

What are the implications of this research for the debate about sex differences in everyday human behaviour? This was very much in the news in 2010 with the publication of Cordelia Fine’s book Delusions of Gender, which was reviewed in The Psychologist, with a reply by Simon Baron-Cohen. Fine focused on two key issues: first, she questioned the standards of evidence used by those claiming biologically-based sex differences in behaviour, and second she noted how there were powerful cultural factors that affected gender-specific behaviour and that were all too often disregarded by those promoting what she termed ‘neurosexism’. I don’t know the literature well enough to evaluate the first point, but on the second, I would agree with Fine that biological factors do not occur in a vacuum. The evidence I’ve reviewed on genes shows unequivocally that there are sex differences in gene expression, but it does not exclude a role for experience and culture. This is nicely illustrated by the research of Michael Meaney and his colleagues demonstrating that gene expression in rats and mice can be influenced by maternal licking of their offspring, and that this in turn may differ for male and female pups!  Genes are complex and fascinating in their effects, but they are not destiny.


Further reading
Davies, W., & Wilkinson, L. S. (2006). It is not all hormones: Alternative explanations for sexual differentiation of the brain. Brain Research, 1126, 36-45. doi: 10.1016/j.brainres.2006.09.105.
Gould, L. (1996). Cats are not peas: a calico history of genetics: Copernicus.
Lemos, B., Branco, A. T., & Hartl, D. L. (2010). Epigenetic effects of polymorphic Y chromosomes modulate chromatin components, immune response, and sexual conflict. Proceedings of the National Academy of Sciences, 107(36), 15826-15831.doi/10.1073/pnas.1010383107.
Skaletsky, H., Kuroda-Kawaguchi, T., Minx, P. J., Cordum, H. S., Hillier, L., Brown, L. G., et al. (2003). The male-specific region of the human Y chromosome is a mosaic of discrete sequence classes. Nature, 423(6942), 825-837.doi: 10.1038/nature01722

Wijchers PJ, & Festenstein RJ (2011). Epigenetic regulation of autosomal gene expression by sex chromosomes. Trends in genetics : TIG, 27 (4), 132-40 PMID: 21334089





Thursday, 21 April 2011

The burqa ban: what's a liberal response?



© www.CartoonStock.com

Most liberal-thinking people regard it as a given that we should respect and tolerate the beliefs of others, even if we don’t share them. However, this can land us in difficulties if the people who hold those beliefs don’t reciprocate. This seems to me at the heart of the debate on the burqa ban.

I vividly remember the time when Salman Rushdie was receiving death threats because of the publication of the Satanic Verses. I had fully anticipated that public figures in the UK would rally around him with robust support for freedom of speech. In fact, the support was muted. Although in part this was because of fear - as noted by Christopher Hitchens, - there were others who clearly felt a tension between such support and a need to empathise with the offence caused by the book. Roy Hattersley, for instance, recommended against publication of a paperback version of the book. And the Chief Rabbi wrote to The Times (4 March 1989) that 'the book should not have been published' because of the need to respect and 'generate respect' for other people's religious beliefs. Subsequently, when Rushdie was offered a knighthood, it was that staunch liberal, Shirley Williams, who emphasised the offence to Muslims, leaving it to Christopher Hitchens and the right-wing Boris Johnston to defend freedom of speech. So an asymmetric relationship was validated: I can’t criticise you because it might offend you, but you can not only criticise what I say, but also insist that I don’t say it at all, on pain of death.

Ayaan Hirsi Ali emphasised a similar trend in her memoir Infidel: liberals, she argued, were reluctant to take action against practices such as forced marriages and genital mutilation, because they did not like to be seen to be criticising another culture. Never mind that the culture was inflicting physical and mental damage on its women. This kind of logic was taken to an extreme by Germaine Greer, who argued that attempts to outlaw genital mutilation were an ‘attack on cultural identity’.

My own views on the matter are quite simple. I will tolerate the views of others so long as they tolerate me. I will respect their cultural identity so long as it does not discriminate against others on the basis of sex, ethnicity, or sexual identity. But I expect my cultural identity and beliefs to be correspondingly respected.

So where does that leave the burqa?

Some liberals adopt the easy argument and say that the burqa is a symbol of oppression, and should therefore be banned. There’s no doubt that the burqa has been used to oppress women, most notably by the Taliban. But it is an oversimplification to argue that all women who wear a burqa are oppressed. There are some (including the young Ayaan Ali Hirsi) who choose to wear it. Yes, that choice is bound to be influenced by the attitudes of those around her, but that is equally true of any woman’s choice of attire, whether it be stiletto heels and a mini-skirt or an all-encompassing robe.

So if it comes down to a woman’s right to choose what to wear, what’s the problem? The issue was mocked on Radio 4’s News Quiz last week, as the participants called for bans on other offensive items of clothing, such as socks with sandals or culottes. Andy Hamilton described the French attitude as: “We will force them to be liberated and if they refuse we will put them in prison”.

But the burqa is different from other clothing choices in two important ways. First, it interferes with communication. Liberal-minded people in the UK have no problem with others wearing symbols of their religion such as a turban or headscarf. The real problem is that in face-to-face interactions, wearing a burqa is at best discourteous, and at worst threatening. It creates an unequal relationship, when you can’t even verify the identity of the person you are interacting with, let alone read facial cues. There’s a vast literature in social psychology looking at how nonverbal cues are important in interpersonal communication (e.g. Knapp & Hall, 2009) . We use them to judge another person’s attitude, honesty, boredom and engagement, for a start. If one person has access to these cues and the other does not, that creates an asymmetry in the interaction. Just as British people must learn to respect the culture of others by not wearing skimpy clothing in Arab countries, Islamic women should respect cultural expectations that it is important to see the face of someone we interact with in person.

The second problem with the burqa is the rationale behind its adoption. As many in the Islamic community have emphasised, the burqa is not mandatory attire for a religious woman. However, Islamic women are required to dress modestly, and some interpret this as requiring total cover-up. I suspect that for some women this is an extreme reaction to our highly-sexualised Western society. But what message does the burqa give to men? I think it is offensive insofar as it implies that they are sexual predators whose lust may be inflamed by the sight of a woman’s face. I liked the response from a participant at the UN Human Rights Council, who suggested that rather than expecting women to cover up, men should stay indoors until they learned some self-control.

So if a woman is to wear a burqa, she should be aware of the impact on others: her choice will appear discourteous to many people, especially men. Qanta Ahmed, who describes herself as a moderate Muslim, makes a related point, noting that, far from giving an impression of modesty, part of that impact is to draw attention to oneself, make others feel threatened and increase hostility to Islam.

I don’t agree with Ahmed that the burqa should be banned; this will simply increase intolerance on both sides. I’m pleased that the British government is showing no signs of going down the same route as the French.  Nevertheless, while I don’t think this is an issue that should be dealt with by law, I do think it is reasonable to exert social pressure. A woman has a right to cover up if she wishes, but she should be aware that this is regarded as culturally inappropriate in many situations in Western society. Employers have a right to expect their staff to have a sense of what is appropriate dress for a job; in many situations, a burqa is no more appropriate than hot pants or a crop top.

So in sum, I would defend the right of a woman to wear a burqa if she chooses. But I would also defend the right of someone like MP Philip Hollobone to refuse to meet with a constituent unless she reveal her face, and for an employer to require more appropriate clothing for someone interacting with the public. To do otherwise is to treat the burqa-wearing woman as someone who has rights but no responsibilities.

Wednesday, 13 April 2011

A short nerdy post about use of percentiles in statistical analyses

Results from psychological tests can be expressed in various ways. Percentiles are a popular format in clinical reports, because they can be explained to non-experts fairly easily, in terms of the percentage of the population that would be expected to get a score of this level or below. So if your score is at the 10th percentile, only 10% of the population would be expected to score this low.
The other format that is commonly used in reporting test scores is the standard score or scaled score. This represents how many standard deviations a score is above or below the population mean. The simplest version is the z-score, obtained by the formula:
  (X-M)/S
where X is the obtained score, M is population mean, and S is the population standard deviation.
In clinical tests, z-scores are often transformed to a different scale, e.g. mean of 100 and SD 15 in the case of most IQ tests. This is done just by multiplying the z-score by the SD and adding the new mean.  So a z-score of -.33 becomes a scaled score of  (-.33 x 15)+100 = 95.
The important point to note is that all of these different methods of reporting scores are just transformations of one another. If you want to turn a z-score into a percentile, you can do so with the Excel function:
100*NORMDIST(A1,0,1,1)
where A1 is the address of the value you want to convert.
The second value in this expression is the mean and the third is the SD, so if you want to convert a scaled score with mean 100 and SD 15 into a percentile, the function is:
100*NORMDIST(A1,100,5,1)
The normsdist function returns a cumulative proportion, so it’s multiplied by 100 to give a percentage.
You can work the other way round using the NORMSINV function, which turns a proportion into a z-score. So if you have a percentile in cell A1, then you get a z-score with:
=NORMSINV(A1/100)

If all this Excel stuff gives you a headache, you can ignore it, so long as you get the message that z-scores, scaled scores and percentiles are all different ways of representing the same information.
They are NOT equivalent, however, in their distributions. Percentiles aren’t suitable as input to statistical procedures that assume normality, such as Anova and t-tests. They should always be converted to z-scores or other scaled scores.
This can be simply illustrated. If you are into Excel, you can generate your own data to make the point - otherwise you can just look at the output from the data I have generated.
Let’s simulate data from two groups, each of 50 participants. Assume the data are reading test scores, and that group 1 has reading difficulties and group 0 hasn’t.  For group 0 I will just generate a random normal distribution of scores with mean 0 and SD 1, by typing this function in each of 50 cells:
=NORMSINV(RAND())
For group 1, I use the same formula, but subtract 0.4 from each score:
=NORMSINV(RAND())-0.4
I pasted my simulated data into SPSS, as it makes it a bit easier to generate relevant statistical output. So for each of 100 simulated subjects, I have a column denoting their group (0 or 1), a column with their z-score, and a column with their percentile score.
Here’s what you get if you do a t-test (you’ll get different values if you generated your own data as the random process is different each time - but it should show the same pattern):
So why, if the numbers are just transforms of each other, are the results different?
The answer lies in the distribution of data. If you take percentiles, you transform a normal distribution into a rectangular one, as can be seen if you plot the histograms.
 
(That hole in the middle of the percentile distribution is just a fluke in the particular dataset I generated). Another way to think about it is to consider the size of difference between two points in the distribution. In terms of z-scores, the difference between the 1st and 10th percentile is 2.32-1.28 = 1.04, and the difference between the 41st and 50th percentile is .23. But on the percentile scale, these differences are treated as equivalent. In effect, the percentile transformation stretches out the points in the middle of the scale and gives them more weight than they should have.
So percentiles are a good way of communicating test scores of individuals, but a bad choice if you are doing statistical analyses of group data.




Saturday, 9 April 2011

Special educational needs: will they be met by the new Green Paper proposals?

Many children with disabilities in the UK are not getting the help they need. The campaigning group Whizz-Kidz states: “There are around 70,000 disabled children in the UK who are waiting to get the wheelchair that suits them best. They can wait months, sometimes years”.  And the situation for children with speech, language and communication needs is equally stark. According to the Bercow report, “The current system is characterised by high variability and a lack of equity. (It) is routinely described by families as a 'postcode lottery', particularly in the context of their access to speech and language therapy.”

A recent Green Paper on Special Educational Needs (SEN) released by the Department of Education takes such concerns on board, stating:
The reforms we set out in this Green Paper aim to provide families with confidence in, and greater control over, the services that they use and receive. For too many parents, their expectations that services will provide comprehensive packages of support that are tailored to the specific needs of their child and their family are not matched by their experiences, just as frontline professionals too often are hampered and frustrated by excessively bureaucratic processes and complex funding systems.and has a wide range of recommendations.” (point 29). This all sounds excellent, so what exactly is proposed?

The two examples given above, of provision of wheelchairs and speech and language therapy are given particular focus. Specifically, the recommendation is for “the option of a personal budget by 2014 for all families with children with a statement of SEN or a new Education, Health and Care Plan” (point 6). There is an emphasis on bringing together the different agencies that are concerned with children who have complex educational needs, so that educational, medical and social agencies work together. So far so good.

Delving deeper into the document, we find the more specific statement:   
“2.41 We have consulted on the introduction of patient choice of any willing provider that meets NHS standards and price for most NHS-funded services by 2013-14. This is likely to apply to many community health services. It will give families choice, where appropriate, from a range of providers who are qualified to provide safe, high quality care and treatment, and select the one that best meets their needs. It will mean that good providers that offer innovative and responsive services are able to grow.”

Note the use of the word “patient” here; although the document talks about “educational, medical and social agencies” working together, the description of the personal budget appears to relate just to health needs.

Not surprisingly, then, the proposed solution has strong parallels with current policies on provision of healthcare, with a focus on outsourcing to private providers. The personal budget is specifically mentioned in relation to provision of such facilities as wheelchairs or speech and language therapy services, which are currently provided via the National Health Service. One can see that any policy that ensures that children get what they need in a timely fashion is to be welcomed, and there is ample evidence that the current system has not always provided this.

The key question, of course, is whether the personal budget will be adequate to give children what they need. All too often, governments have dressed up healthcare policies as providing more “choice”, when in fact they are designed to save money. It is inconceivable in the current economic climate that any extra money will be available for disability and SEN. So everything hinges on whether the personal budget will be sufficient to cover a child’s needs. No doubt the expectation is that competition between private providers will drive down costs so it will be possible to “do more with less”. I hope this works, but I'm not optimistic.

But what about the educational and social aspects of provision? I’m particularly interested in whether the “patient choice” model will be extended from the medical to the educational sphere. Here there is a potential problem. For medical provision, it is argued that NHS standards must be met. One would hope, then, that there will be vetos on spending the personal budget on such interventions as “acupuncture miracle cure” or stem cell treatment for cerebral palsy.  But in the field of special education it’s not clear what standards would need to be met.

Evidence-based education is still in its infancy, and in mainstream education there are plenty of instances where government funds have been spent on educational programmes of dubious or unproven effectiveness. Ben Goldacre, in his book Bad Science, documented the way in which Brain Gym programmes were introduced in UK schools, despite being full of ludicrous pseudoscience. Charlie Brooker’s account of this is also worth a read. Though we may laugh at these educational initiatives, they provide a nice income stream for those who are marketing them. (See also, http://wordpress.mrreid.org/2009/06/01/learning-styles-are-nonsense/)

If the plan is to give families a personal budget to spend on special education, there will be plenty of companies who will see this as a fantastic business opportunity. Some may be providing beneficial services, but there is a real risk that commercial companies will be rewarded with government funds for interventions of dubious or unknown value. Suppose your child has severe reading difficulties, language comprehension problems, or autistic features, and the classroom teacher seems at a loss to know how to help. It is easy to envisage a situation where private companies could offer attractive-sounding interventions, in anticipation that the “personal budget” could be used to support these. There are already numerous cases of such fringe interventions, but currently any parent who wants to try them has to find their own funding. In most cases, the only evidence for efficacy is anecdotal. In a few, claims of scientific support are made, but usually when investigated, the evidence proves to be weak. Randomized controlled trials are very rare in the field of education, and where these have been applied to interventions for children’s learning and educational difficulties, results have typically been much less impressive than when uncontrolled studies are done.

Clearly, if we demanded that any educational approach used in schools had to be demonstrated to be effective to a high standard of evidence, the school system would grind to a halt. Education has never been required to meet the standards of evidence seen in medicine. Furthermore, no innovations would ever occur. I’m a great fan of changing this system to one of “evidence-based education”, but I am realistic enough to realise it is not going to happen overnight. And if we do take an evidence-based approach, we need to consider carefully the way in which we measure children’s outcomes: it could be a mistake to have a narrow focus on educational attainment that does not consider other aspects of well-being. But we need to address these issues urgently if we are going to give “any willing provider” the opportunity to sell services for children with special needs.

As an example of the potential pitfalls, it is worth reading a 2008 report of the Enterprise and Learning Committee of the Welsh Assembly   It is noteworthy that some of those advising the committee had vested interests in the programmes under discussion: Prof David Reynolds had financial interests in the Dore programme, and evidence for efficacy of FastForword was provided by those who sold the programme, and a professor whose institution had made $5.5 million from royalties.  The committee were apparently unaware of independent evaluations of these programmes that gave a much less positive picture, see: http://tinyurl.com/3q4jen9  and a recent meta-analysis of FastForword, which includes studies published prior to 2008 (not to be confused with the Fast Forward wheelchair campaign).

I am not saying that private companies should be excluded from providing special education. Potentially, they have much to offer: being freed from bureaucracy of state-based organisations they have potential to develop new approaches, or to deliver traditional services in an exemplary fashion. But we need to have stringent standards in place when evaluating “willing providers” and ensure there is no conflict of interest in those advising on appropriate interventions. Some providers will see this population as a wonderful commercial opportunity. We need to ensure that limited funds are spent wisely and well.

Of course, this is only germane if a personal budget will be available for educational as well as medical interventions. I wonder whether it will be. The Green Paper is not clear on this, noting merely that “Subject to piloting, this would include funding for education and health support as well as social care”. But even if the piloting supports educational uses for a personal budget, it is not clear which children would have access to this. On the one hand, the Green Paper emphasises the numerous ways in which children currently identified with SEN have poor outcomes. Yet on the other hand it seems to imply that many of those with labels of SEN don’t have genuine problems: “Previous measures of school performance created perverse incentives to over-identify children as having SEN. There is compelling evidence that these labels of SEN have perpetuated a culture of low expectations and have not led to the right support being put in place.” (point 22). And “we intend to tackle the practice of over-identification by replacing the current SEN identification levels of School Action and School Action Plus with a new single school-based SEN category for children whose needs exceed what is normally available in schools; revising statutory guidance on SEN identification to make it clearer for professionals; and supporting the best schools to share their practices." (point 24, my emphasis). Finally, “A new single early years setting- and school-based category of SEN will build on our fundamental reforms to education which place sharper accountability on schools to make sure that every child fulfils his or her potential.” (point 5, my emphasis). 

The Green Paper sounds full of good intentions, but I’m cynical. Cutting through the fine language I see a pincer movement to cut costs of children with disabilities and SEN: first, by radically reducing the number of children who will be deemed to need special provisions, and second, by passing responsibility for the remainder over to the marketplace. I fervently hope the recommendations will do good in overcoming the obstacles currently faced by families in obtaining necessary equipment and services, but I have two worries. First, that the profit motive of those in the marketplace might conflict with the child’s best interests, and second that the net impact for children with hidden disabilities will be to reduce provision and then blame teachers for children’s educational failure.

The Green Paper is a consultation document; I'd encourage all readers with an interest in this area to read the document and give your own views: you have until 30th June 2011 to respond.

Sunday, 20 March 2011

The expansion of research regulators: an evolutionary perspective

Reading about evolution has made me think about why some professions grow and thrive while others die out. I'm intrigued by the expansion in numbers of people regulating the activities of researchers. How have we got to a position where the Academy of Medical Sciences concludes: “A complex and bureaucratic regulatory environment is stifling health research in the UK”?

Consider the situation in the 1970s. If you wanted to do a piece of research you did it, no questions asked. But bad things can happen if you let people do just what they want. There are terrible examples of studies where research participants were infected, hurt or humiliated without realising what was happening or giving their consent. For examples, see Rebecca Skloot’s book, ‘The Immortal Life of Henrietta Lacks’ and Dominic Streatfeild’s ‘Brainwash: the Secret History of Mind Control’. The solution was to create a body of people, the regulators, who would scrutinise research and make sure it was ethical. Despite the regulation, every few years something bad still happened. The regulators responded by increasing their numbers and adding more regulations. In general, I’ve avoided doing studies that require me to go through an NHS ethics committee because the process is so long-winded and bureaucratic that it saps all my enthusiasm and takes up time I’d rather spend doing research. Our University ethics committee can approve studies that don’t involve patients and operates a much less complex system. Recently, though, I badly wanted to do a study involving NHS patients, and decided to grit my teeth and go through the process. It’s taken literally weeks of form-filling, and what amazed me was the sheer number of regulators I dealt with in the course of applying for approval. Then there was one set of people from the research ethics committee (REC), another set from R&D, yet more from the Comprehensive Clinical Research Network - in fact several sets of those depending on whether you were concerned with local or regional matters. There are people whose job it is to book your application in to a REC via a centralised system, and others whose job it is to do the same thing at local level when the first fail to find you a slot. These people were typically very helpful, but that isn’t the point. Why are there so many people whose sole function in life is the ethical scrutiny of researchers? Why are there so many forms to fill in that a recent article raised concerns about the environmental impact of paper use by RECs.  And how have we got to the situation described by the Academy of Medical Sciences whereby it takes an average of 621 days from receiving funding to recruiting the first patient for a trial of a cancer drug? 

From an evolutionary perspective, a research regulator is a life form with three very interesting characteristics. First, its numbers explode in response to catastrophic events regardless of how rare that event is. Second, it has few natural predators, so its expansion goes unchecked. Third, regulators multiply like bacteria: they spawn more regulations which require more regulators, so there is a rapid increase in population over time. And these three characteristics derive, I submit, from a basic human tendency to focus on emotionally-engaging events while ignoring their probability.

Catastrophe as a driving force in increasing the number of regulators
When something really terrible happens - someone is badly hurt or upset, or even worse, killed - we empathise with the victim and want to do something to prevent it happening again. All of our attention is taken up by the awfulness of the event, and we ignore the costs inherent in a solution. This kind of thinking is described by Dan Gardner in his book ‘Risk’ as due to System 1, or Gut, as opposed to the more rational System 2, or Head. Gut’s supremacy is such that if someone were to draw attention to rarity of the catastrophe or the costs of the proposed solution, they would be criticised for being heartless. It is this way of reasoning that fosters the dramatic rise in regulators.

Consider the case of Dr Harold Shipman, a general practitioner in Greater Manchester, who in 2000 was found guilty of murdering 15 of his patients. According to Wikipedia, his was one of the most prolific known serial killers in global history with 215 murders being positively ascribed to him, although the real number is likely to be higher than this. He had no obvious motive and did not appear mentally ill to his colleagues or patients. This case led to the Shipman enquiry, led by Dame Janet Smith.  It was discovered that Shipman had been sent a warning letter by the GMC but allowed to return to practice after a conviction for dishonestly obtaining pethidine in 1976.  The enquiry judged, however, that if a harsher punishment had been given, it would not have prevented Shipman from becoming a serial killer. Nevertheless, the committee called for a database to be established containing information about all doctors in the NHS, including disciplinary records, which both patients and NHS bodies could access.  They also supported a system of revalidation, whereby doctors would undergo regular checks of their competency to practise.  There is no indication that anyone ever discussed the probability of another Harold Shipman occurring. I’m sure there are many doctors who are incompetent, have massive personal problems, and there are no doubt a few who feel like murdering their patients from time to time. But I find it hard to believe that we need to set up regulations to scrutinise apparently sane doctors to prevent them from murdering their patients in cold blood.  Nevertheless, in the interests of ‘this must never happen again’ it’s recommended that a whole new posse of regulators be created to check family doctors, whose compliance will no doubt cost time that could be spent with their patients. I suspect one day there will be another doctor who does something really terrible, but I doubt these regulations would prevent this.

A much less dramatic but pertinent example was described on Jenny Rohn’s blog. I recommend you read her account of the regulations produced by her research funder that require that staff in the laboratory wear safety glasses at all times. Jenny, a woman after my own heart, took the trouble to get to the bottom of why this regulation had been introduced, and found there had been a small number of accidents, which could have been prevented if the scientist had taken commonsense precautions and worn safety glasses while performing specific hazardous procedures. A reminder to staff to do this should have been sufficient. Instead, a regulation has been introduced which costs time and money.

Regulators have no natural predators
Once regulation is established, it is remarkably difficult to remove it. This is largely a consequence of the same human tendency as discussed above: the attentional focus on catastrophe. Anyone who argues against regulation will be seen as being so cold-hearted or cavalier as to not care about the catastrophe that led to regulation being set up.

A key point here is that individual regulations often appear trivial - especially when considered in relation to the catastrophes they are designed to avert. Filling in a form, or going to an opticians is tedious, but it seems curmudgeonly to complain if someone’s life or sight can be saved. However, there are expenses in both time and money, and these can become substantial if large numbers of people are required to adhere to regulations and to administer them. We do need to consider carefully whether the measures that are put in place are effective and proportionate.

Consider another example. On 4th August 2002, 10-year-olds Holly Wells and Jessica Chapman were murdered by their school caretaker Ian Huntley in the village of Soham, Cambridgeshire. Huntley had a string of previous allegations about sexual interest in young girls when previously in the North East of England, as well as a burglary charge, but only the burglary charge was placed on the police national computer, and even this was not picked up by the routine checks that the school did, because Huntley had changed his name.  After this case, there has been massive tightening up of police checks for people who work with children. If you plan to work with children or young people, you need a Criminal Records Bureau (CRB) check.  My researcher team works in schools and we all have CRB checks. Recently, though, we’ve found some head teachers will want a new CRB check, just for their school, even if you have recently obtained one. And the regulations have been extended to individuals such as children’s authors who make occasional visits to schools.  Everyone is clearly very nervous about letting unvetted adults come into contact with children. But does it work? On 1st October 2009, Plymouth nursery worker Vanessa George admitted 13 charges of sexual abuse of children and making and distributing indecent images of children. She had completed a qualification in child care and passed a Criminal Records Bureau police check to allow her to work with younger children. 

I am aware that if I query the usefulness of the CRB check procedures it will look as if I am placing my own personal inconvenience above the welfare of vulnerable children. Regulation is tedious, and sometimes costly, but what monster would refuse to fill in a form or pay a few pounds in order to prevent a child being murdered? I can assure readers that I feel every bit as much rage and grief as anyone else every time I see that photo of Holly and Jessica that is so often reproduced in the media. If something can be done to stop children getting murdered or molested, I would be the first to endorse it. I just query whether this massive bureaucratic exercise is a cost-effective solution, as compared, say, with using resources to teach children how to identify and respond to adults who behave inappropriately.

Ultimately, the only thing that could lead to a mass extinction of regulators would be if government were to decide that the regulation was too expensive. However, in general, governments are nervous of deregulation because it will upset people who see regulation as the path to preventing another catastrophe. I disagree. It’s my belief that we can never control life so that there are no catastrophes. So from time to time bad things will happen. Every time they do, more regulators are created, but none are ever removed. Their inexorable rise seems inevitable. But as if this were not enough, there is an additional process at work.

Regulators generate more regulators
In the field of ethical scrutiny of research, a major shake-up was spawned by one rare event, the discovery that a pathologist at Alder Hey Children’s Hospital had stored body organs of deceased children without their parents’ knowledge or consent.  This lead to an explosion of regulation, and a new legal framework for the use of human tissue. But perhaps more surprisingly, it was accompanied by a broadening of the remit of research regulation to apply not just to medical research but to all research involving human participants. I suggest that a driving force here is the regulator mindset. Once you have been set up to prevent catastrophic events, you don’t just focus on the original catastrophe that started the ball rolling, you start trying to anticipate catastrophes, so you can set regulations in place to prevent them. This inevitably generates huge amounts of additional regulation. I thoroughly agree with the idea that it is better to anticipate problems than deal with their consequences, but the difficulty here is that the potential catastrophe absorbs all one’s attention and once again its probability is never considered. 

To take an example, some years ago I was part of a research group who wanted to recruit from a local maternity hospital; we simply wanted to sign up mothers who might potentially be interested in taking part in research when their children was 12 to 36 months of age. At that point, they would be contacted and invited to take part, with no obligation to do so. We did not approach mothers of babies who had any medical problems. One member of the ethics committee was concerned at our procedures. It was suggested that before writing to these parents to invite them to take part in a study, we ought to check with the family doctor whether the child had died. Well, of course, I can imagine it would be awful to receive a letter inviting you to involve your child in a study if the child had died. But what proportion of healthy babies die by 2 or 3 years of age? Is the probability of this happening high enough to justify asking family doctors of some hundred children to check their medical records before we contacted them? According to the Office of National Statistics, the mortality rate for children aged 1 to 14 years was 12 deaths per 100,000 in 2009. Deaths in children under one year of age were more common at 4.5 per 1,000 live births, but most of these were babies who would not have been recruited to our study because they were severely ill in the first week of life, and/or had very low birthweight. I am again uncomfortably aware that I will seem heartless in arguing against adopting a measure to avoid the real but rare possibility of upsetting a bereaved parent. But against this hypothetical risk we need to balance the 100 patients that won’t get seen by their family doctor in the 10 minutes it takes them to locate and check the medical records and reply to the researcher.

In case you imagine such scenarios are unique to the UK, let me give one more example, from the USA. A colleague who does brain-scanning studies of children with developmental disorders tells me she was required to conduct a pregnancy test with any girl aged 9 years or over who wished to participate in the study. This is particularly striking because it is protecting against a conjunction of two very rare possibilities: (a) that a 9-year-old girl who volunteers for a research study might be pregnant and (b) that a brain scan of the mother's head might damage a foetus.

When regulators get together with lawyers, there’s a catalytic reaction, because lawyers are even better than regulators at thinking of things that need regulating. They go beyond defending us from catastrophes and disasters to protecting against things that have the potential to upset a few people. I was interested to read in the report by the Academy of Medical Sciences that the NHS Litigation Authority had never received a claim relating to research, yet the lawyers insist we put paragraphs in our information sheets about risk, indemnity and how to make a complaint. They’ve also had a major success defending people’s rights not to have their medical records scrutinised by anyone outside the clinical care team. The problem is that if you want to do a medical research study you need to identify suitable people to take part, and that means looking at their medical records. If you insist, as current regulation requires, that no-one outside the clinical care team can look at records, you have two stark options. Either the clinical care team have to spend time not caring for patients, but trawling through records, or the research cannot get done. Richard Doll was one of the first to speak out against this kind of regulation, which makes most epidemiological studies impossible to do.  This is one point where the report by the Academy of Medical Sciences has a recommendation to allow bona fide researchers to screen medical records.

Another factor leading to multiplication of regulators is an attitude of trusting no-one. Having required researchers to give a detailed account of what their research involves, down to the last comma in an information sheet, they need squads of highly trained people to scrutinise the forms to identify possible problems. And if this is not enough, they then introduce a further stage of monitoring the research. The implication is that the researchers can’t be trusted. Unless they write regular reports to the regulators on the progress of the research, they are likely to go off the rails. Even this is not enough: the regulators also have power to visit researchers to ensure they are doing what they said they would do. This of course all creates more jobs for the regulators. Nobody ever asks whether the money might be better spent on, for instance, doing research.

How can we retrieve the situation?
Having seen the increasing drive for more and more regulation during my lifetime, I am alarmed at its unstoppable progress. I’ve focused here on those aspects that have impinged on my life as a researcher, but the trend for ever more regulation appears to infest many other areas of life. I have two suggestions for how to improve matters:

a) Before any new regulation is introduced, there should be a cold-blooded cost-benefit analysis that considers (i) the severity of the adverse event that the regulation is designed to avert; (ii) the probability of the adverse event;  (iii) the likely impact of the regulation in reducing that probability; (iv) the cost of the regulation both in terms of the salaries of people who implement it, and the time and other costs to those affected by it. I use the word cold-blooded deliberately: our normal human instincts don’t lead us to weigh up these different factors rationally. Instead, we focus solely on (i).

b) We should be more imaginative about the type of regulation that is used. For instance, research regulators increasingly play a role in training researchers in ethical conduct of research. Currently one is expected to undertake such training in addition to all the form-filling. But why not treat it more like a driving test? Once trained, researchers could be certified as competent and left to get on with it without having to fill in any forms, and without constant scrutiny and monitoring. It would save huge amounts of everyone’s time and money if we could trust people to behave professionally and treat ethical skills more like driving skills. The regulators could then focus on training researchers and offering advice to those who encountered specific ethical issues in their research. Their role would become advisory rather than policing.

I had intended to write a blogpost documenting the many stages I have gone through on the road to seeking ethics approval for my current study, but that procedure, started in December, is continuing, and I cannot tell when it will end.