Showing posts with label diagnosis. Show all posts
Showing posts with label diagnosis. Show all posts

Monday, 20 October 2025

A LEAP into the future, or off a cliff: Wellcome LEAP's new $50M program

A few days ago, I saw this post on LinkedIn:

How does the gut microbiome shape early brain development? That’s what FORM, a new $50 million programme from Wellcome Leap, aims to answer. Critically, it wants to identify the role of the microbiome in autism and other neurological disorders. Applicants from universities, companies and non-profits are invited to submit project proposals by 14 November. 

The full programme announcement for FORM (Foundations of a Resilient Microbiome) that you can download here hops around citing various references that indicate the microbiome is important for early development and can be influenced by factors such as antibiotics. So far, so uncontroversial. But then the topic of autism is introduced.

First we hear that autism diagnoses have increased. Then, a section devoid of references states:

 Many have attributed this increase to expanded surveillance, broadening of diagnostic categories (to include milder autism-related difficulties), or increased public awareness. While all are true, the significance of the increase suggests other rising risk factors may also be contributing. 

We then are told that this can't be due to genes because they are pretty stable in populations, and so we seem led remorselessly to the conclusion that it must be an environmental factor, and what better culprit could there be than the microbiome.

If I was going to predicate a $50 million research program on that premise, I'd do a bit more research into those studies on the increase in diagnosis. Diagnostic criteria have changed radically, so children who would have in the past had other diagnoses, or no diagnosis, are now encompassed within autism. Furthermore, there is wider understanding of autism, and a diagnosis can bring with it educational support, which can be a reason why parents will seek a diagnosis. Here's a simple explainer that I wrote in 2012, and subsequent studies by Lundström et al (2015)Cardinal et al (2016) and Zeidan et al (2022).

The fragments of supportive evidence that are provided for the autism/microbiome link seem cherrypicked and are not impressive. At least one claim seems just plain wrong: "Babies exposed to antibiotics in the first 6 months of life may be twice as likely to develop ASD as those exposed later." The cited paper by Azad et al (2016) doesn't mention autism and I when I searched for another source, I found a solid-looking study claiming no association. Other cited results are the kinds of findings that you get if you test for so many associations that some are bound to come up by chance. The handful of animal model studies that are cited have been criticised on methodological grounds.

Things go more seriously off the rails when specific quantitative goals are set for the program: 

Autism currently affects about 3.2% of children. To identify what proportion of these cases may be attributable to gut microbiome dysfunction, we will need objective biomarkers that can detect the dysfunction with high accuracy (balanced accuracy >90%). Establishing this will require a large cohort — likely more than 15,000 children — to ensure statistical power. With that sample size, we can reliably estimate whether microbiome dysfunction accounts for as much as 50% of ASD cases (around 1.6% of all children) or as little as 10% (about 0.3%).

Three things about this:

  1. In a footnote it is noted that "Severe autism with an established genetic origin (about 10–20%) and mild/ moderate ASD (~40%) fall outside the scope of this program" - these estimates don't seem to take that into account. And it's not clear if the databases that will be used for the analysis actually allow one to distinguish autism subtypes.
  2. You can't establish causality from observational data. As has been shown by Yap et al (2021), the microbiome is influenced by specific dietary preferences of autistic children.
  3. It is assumed that the lower bound is that microbiome dysfunction explains 10% of cases. This shows remarkable commitment to a causal hypothesis that has no solid evidence: a realistic lower bound would be zero.

As regards the plans in "Thrust 2" to have a diagnostic set of biomarkers that will predict severe autism-related difficulties, there are so many issues here, that I recommend reading previous blogposts I wrote on this topic, here, and here.  In brief, screening is only effective if there is a strong association between biomarkers and outcomes and if the biomarker measures are stable. Even if those conditions are met, if the base rate of the condition (autism) is low, you will be overwhelmed with false positives.

I was surprised that a reputable funding body was associated with this program, so I wanted to find out more about Wellcome LEAP.  They are a U.S.-based non-profit organization founded by the Wellcome Trust that:

builds bold, unconventional programs, and funds them at scale. Programs optimized to deliver breakthroughs in human health over 5 – 10 years and demonstrate seemingly impossible results on seemingly impossible timelines.

No doubt I'm too conventional, but I'm nervous of "seemingly impossible" things. If they are claimed, I am suspicious, especially if this occurs on "seemingly impossible timelines".

Reading on, I can see lots of things to like about LEAP. The idea is to cut time spent in bureaucratic processes of setting up grants, and to bring together networks of researchers from different institutions and different disciplines who can work together to solve problems that involve large and complex datasets, and check generalisability of findings. This kind of international collaborative approach that involves diverse populations is a definite bonus of LEAP.

The worrying bit was the emphasis on speed - especially since this was coupled with an expectation that the results would be commercialised.

This is clarified here:

Wellcome Leap anticipates that it will normally further our mission (and the organization’s mission) to commercialize the results of Wellcome Leap-funded research. If we determine that a Performer’s organization is not making appropriate efforts to further commercialization, either itself or through a third party (e.g., a licensee), we have the option to request a meeting and a remedial commercialization plan to address the issue.

Given that I have in the past written in favour of slow science, it's perhaps not surprising that this funding model doesn't appeal to me. The thing is that not all delays in scientific progress are down to bureaucracy or timidity. A major obstacle to progress is time wasted trying to build on prior research findings that prove to be illusory.

Science should be cumulative, which means we should be able to proceed with confidence and assume that published research is robust. The likelihood of this being the case is low if we are dealing with complex multidimensional systems, where the temptation is to just hunt around until we find something that looks exciting, embellish it with a few statistical credentials and claim we have a novel result.  As Chin (2025) has argued, when put under pressure to deliver speedy results, scientists may be forced into a position of hyping their findings, cutting corners, and reporting only favourable results. 

I was frankly dismayed to read that in her prior role as Program Director for the Wellcome Leap “First 1000 Days” (1kD) initiative, the Program Director of FORM:

led the delivery of multiple new product opportunities to improve cognitive development in the first 1000 days of a child’s life — including a breakthrough microbiome-directed diagnostic and therapeutic now positioned for commercialization

The 1kD initiative was funded in summer of 2021. So it seems that in four years a microbiome-directed "diagnostic and therapeutic" has been developed and validated. I'm afraid that without solid published evidence, this looks like NeuroPointDX all over again. 

Wellcome LEAP programs are focused around "What If" questions. My question is "What if there is no association between autism and microbiome dysfunction?" All three "thrusts" of this program depend on there being an association. My suggestion is to set aside 1% of the funds for this program for pre-registered replications of the studies that are cited as foundational for the research. I know the idea is to fund high-risk, unconventional research, but a lot of time and money could be saved by checking that the foundations are solid before building an edifice on this premise. 

 Comments on this blog are moderated so may take time to appear.  In general, anonymous comments are not approved. 

Tuesday, 6 December 2022

Biomarkers to screen for autism (again)


Diagnosis of autism from biomarkers is a holy grail for biomedical researchers. The days when it was thought we would find “the autism gene” are long gone, and it’s clear that both the biology and the psychology of autism is highly complex and heterogeneous. One approach is to search for individual genes where mutations are more likely in those with autism. Another is to address the complexity head-on by looking for combinations of biomarkers that could predict who has autism.  The latter approach is adopted in a paper by Bao et al (2022) who claimed that an ensemble of gene expression measures taken from blood samples could accurately predict which toddlers were autistic (ASD) and which were typically-developing (TD). An anonymous commenter on PubPeer queried whether the method was as robust as the authors claimed, arguing that there was evidence for “overfitting”. I was asked for my thoughts by a journalist, and they were complicated enough to merit a blogpost.  The bottom line is that there are reasons to be cautious about the conclusion of the authors that they have developed “an innovative and accurate ASD gene expression classifier”.

 

Some of the points I raise here applied to a previous biomarker study that I blogged about in 2019. These are general issues about the mismatch between what is done in typical studies in this area and what is needed for a clinically useful screening test.

 

Base rates

Consider first how a screening test might be used. One possibility is that there might be a move towards universal screening, allowing early diagnosis that might help ensure intervention starts young.  But for effective screening in that context, you need extremely high diagnostic accuracy, and accuracy depends on the frequency of autism in the population.  I discussed this back in 2010. The levels of accurate classification reported by Bao et al would be of no use for population screening because there would be an extremely high rate of false positives, given that most children don’t have autism.

 

Diagnostic specificity

But, you may say, we aren’t talking about universal screening.  The test might be particularly useful for those who either (a) already have an older child with autism, or (b) are concerned about their child’s development.  Here the probability of a positive autism diagnosis is higher than in the general population.  However, if that’s what we are interested in, then we need a different comparison group – not typically-developing toddlers, but unaffected siblings of children with autism, and/or children with other neurodevelopmental disorders.   

When I had a look at the code that the authors deposited for data analysis, it implied that they did have data on children with more general developmental delays, and sibs of those with autism, but they are not reported in this paper. 

 

The analyses done by the researchers are extremely complex and time-consuming, and it is understandable that they may prefer to start out with the clearest case of comparing autism with typically-developing children. But the acid test of the suitability of the classifier for clinical use would be a demonstration that it could distinguish children with autism from unaffected siblings, and from nonautistic children with intellectual disability.

 

Reliability of measures

If you run a diagnostic test, an obvious question is whether you’d get the same result on a second test run.  With biological and psychological measures the answer is almost always no, but the key issue for a screener is just how much change there is. Gene expression levels could vary from occasion to occasion depending on time of day or what you’d eaten – I have no idea how important this might be, but it's not possible to evaluate in this paper, where measures come from a single blood sample. My personal view is that the whole field of biomedical research needs to wake up to the importance of reliability of measurement so that researchers don’t waste time exploring the predictive power of measures that may be too unreliable to be useful.  Information about stability of measures over time is a basic requirement for any diagnostic measure.

 

A related issue concerns comparability of procedures for autism and TD groups. Were blood samples collected by the same clinicians over the same period and processed in the same lab for these two groups? Were the blood analyses automated and/or done blind? It’s crucial to be confident that minor differences in clinical or lab procedures do not bias results in this kind of study.

 

Overfitting

Overfitting is really just a polite way of saying that the data may be noise. If you run enough analyses, something is bound to look significant, just by chance.  In the first step of the analysis, the researchers ran 42,840 models on “training” data from 93 autistic and 82 TD children and found 1,822 of them performed better than .80 on a measure that reflects diagnostic accuracy (AUC-ROC – which roughly corresponds to proportion correctly classified: .50 is chance, and 1.00 is perfect classification).  So we can see that just over 4% of the models (1822/42840) performed this well.

 

The researchers were aware of the possibility of overfitting, and they addressed it head-on, saying: “To test this, we permuted the sample labels (i.e., ASD and TD) for all subjects in our Training set and ran the pipeline to test all feature engineering and classification methods. Importantly, we tested all 42,840 candidate models and found the median AUC-ROC score was 0.5101 with the 95th CI (0.42–0.65) on the randomized samples. As expected, only rare chance instances of good 'classification' occurred.”  The distribution of scores is shown in Figure 2b. 

 

 


Figure 2b from Bao et al (2022)

 

They then ran a further analysis on a “test set” of 34 autistic and 31 TD children who had been held out of the original analysis, and found that 742 of the 1822 models performed better than .80 in classification. That’s 40% of the tested models.  Assuming I have understood the methods correctly, that does look meaningful and hard to explain just in terms of statistical noise.  In effect, they have run a replication study and found that a substantial subset of the identified models do continue to separate autism and TD groups when new children are considered. The claim is that there is substantial overlap in the models that fall in the right-hand area under the curve for the red and pink distributions.

 

The PubPeer commenter seems concerned that results look too good to be true. In particular, Figure 2b suggests the models perform a bit better in the test set than in the training set. But the figure shows the distribution of scores for all the models (not just the selected models) and, given the small sample sizes, the differences between distributions does not seem large to me. I was more surprised by the relatively tight distribution of AUC-ROC values obtained in the permutation analysis, as I would have anticipated some models would have given high classification accuracy just by chance in a sample of this size.

The researchers went on to present data for the set of models that achieved .8 classification in both training and test sets. This seemed a reasonable approach to me. The PubPeer commenter is correct in arguing that there will be some bias caused by selecting models this way, and that one would expect  less good performance in a completely new sample, but the 2-stage selection of models would seem to ensure there is not "massive overfitting". I think there would be a problem if only 4% of the 1822 selected models had given accurate classification, but the good rate of agreement between the models selected in the training and test samples, coupled with the lack of good models in the permuted data, suggests there is a genuine effect here. 

 

Conclusion

So, in sum, I think that the results can’t just be attributed to overfitting, but I nevertheless have reservations about whether they would be useful for screening for autism.  And one of the first things I’d check if I were the researchers would be the reliability of the diagnostic classification in repeated blood samples taken on different occasions, as that would need to be high for the test to be of clinical use.

 

Note: I'd welcome comments or corrections on this post. Please note, comments are moderated to avoid spam, and so may not appear immediately. If you post a comment and it has not appeared in 24 hr, please email me and I'll ensure it gets posted. 

 PS. See comment from original PubPeer poster attached. 

Also, 8th Dec 2022, I added a further PubPeer comment asking authors to comment on Figure 2B, which does seem odd. 

https://pubpeer.com/publications/B693366B2B51D143C713359F151F7B#4  


PPS. 10th Dec 2022. 

Author Eric Courchesne has responded to several of the points made in this blogpost on Pubpeer: https://pubpeer.com/publications/B693366B2B51D143C713359F151F7B#5  



 


 

 

 

 

 

 

 

 

Saturday, 12 January 2019

NeuroPointDX's blood test for Autism Spectrum Disorder: a critical evaluation

NeuroPointDX (NPDX), a Madison-based biomedical company, is developing blood tests for early diagnosis of Autism Spectrum Disorder (ASD). According to their Facebook page, the NPDX ASD test is available in 45 US states. It does not appear to require FDA approval. On the Payments tab of the website, we learn that the test is currently self-pay (not covered by insurance), but for those who have difficulty meeting the costs, a Payment Plan is available, whereby the test is conducted after a down payment is received, but the results are not disclosed to the referring physician until two further payments have been made.

So what does the test achieve, and what is the evidence behind it?

Claims made for the test
On their website, NPDX describe their test as a 'tool for earlier ASD diagnosis'. Specifically they say:
'It can be difficult to know when to be concerned because kids develop different skills, like walking and talking, at different times. It can be hard to tell if a child is experiencing delayed development that could signal a condition like ASD or is simply developing at a different pace compared to his or her peers...... This is why a biological test, one that’s less susceptible to interpretation, could help doctors diagnose children with ASD at a younger age. The NPDX ASD test was developed for children as young as 18 months old.'
They go on to say:
'In our research of autism spectrum disorder (ASD) and metabolism, we found differences in the metabolic profiles of certain small molecules in the blood of children with ASD. The NPDX ASD test measures a number of molecules in the blood called metabolites and compares them to these metabolic profiles.

The results of our metabolic test provide the ordering physician with information about the child’s metabolism. In some instances, this information may be used to inform more precise treatment. Preliminary research suggests, for example, that adding or removing certain foods or supplements may be beneficial for some of these children. NeuroPointDX is working on further studies to explore this.

The NPDX ASD test can identify about 30% of children with autism spectrum disorder with an increased risk of an ASD diagnosis. This means that three in 10 kids with autism spectrum disorder could receive an earlier diagnosis, get interventions sooner, and potentially receive more precise treatment suggestions from their doctors, based on information about their own metabolism.'
They further state that this is:  'A new approach to thinking about ASD that has been rigorously validated in a large clinical study' and they note that results from their Children’s Autism Metabolome Project (CAMP) study have been 'published in a peer-reviewed, highly-regarded journal, Biological Psychiatry'.

The test is recommended for a child who:
  • Has failed screening for developmental milestones indicating risk for ASD (e.g. M-CHAT, ASQ-3, PEDS, STAT, etc.). 
  • Has a family history such as a sibling diagnosed with ASD. 
  • Has an ASD diagnosis for whom additional metabolic information may provide insight into the child’s condition and therapy.
In September, Xconomy, which reports on biotech developments, ran an interview with Stemina CEO and co-founder Elizabeth Donley, which gives more background, noting that the test is not intended as a general population screen, but rather as a way of identifying specific subtypes among children with developmental delay.

Where are the non-autistic children with developmental delay? 
I looked at the published paper from the CAMP study in Biological Psychiatry.

Given the recommendations made by NPDX, I had expected that the study would involve comparison of children with developmental delay to compare metabolomic profiles in those who did and did not subsequently meet diagnostic criteria for ASD.

However, what I found instead was a study that compared metabolomics in 516 children with a certified diagnosis of ASD and 164 typically-developing children. There was a striking difference between the two groups in 'developmental quotient (DQ)', which is an index of overall developmental level. The mean DQ for the ASD group was 62.8 (SD = 17.8), whereas that of the typically developing comparison group was 100.1 (SD = 16.5). This information can be found in Supplementary Materials Table 3.

It is not possible, using this study design, to use metabolomic results to distinguish children with ASD from other cases of developmental delay. To do that, we'd need a comparison sample of non-autistic children with developmental delay.

The CAMP study is registered on ClinicalTrials.gov, where it is described as follows:
'The purpose of this study is to identify a metabolite signature in blood plasma and/or urine using a panel of biomarker metabolites that differentiate children with autism spectrum disorder (ASD) from children with delayed development (DD) and/or typical development (TD), to develop an algorithm that maximizes sensitivity and specificity of the biomarker profile, and to evaluate the algorithm as a diagnostic tool.' (My emphasis)
The study is also included on the NIH Project Reporter portfolio, where the description includes the following information:
'Stemina seeks funding to enroll 1500 patients in a well-defined clinical study to develop a biomarker-based diagnostic test capable of classifying ASD relative to other developmental delays at greater than 80% accuracy. In addition, we propose to identify metabolic subtypes present within the ASD spectrum that can be used for personalized treatment. The study will include ASD, DD and TD children between 18 and 48 months of age. Inclusion of DD patients is a novel and important aspect of this proposed study from the perspective of a commercially available diagnostic test.' (My emphasis)
So, the authors were aware that it was important to include a group with developmental delay, but they then reported no data on this group. Such children are difficult to recruit, especially for a study involving invasive procedures, and it is not unusual for studies to fail to meet recruitment goals. That is understandable. But it is not understandable that the test should then be described as being useful for diagnosing ASD from within a population with developmental delay, when it has not been validated for that purpose.

Is the test more accurate than behavioural diagnostic tests? 
A puzzling aspect of the NPDX claims is a footnote (marked *) on this webpage:
'Our test looks for certain metabolic imbalances that have been identified through our clinical study to be associated with ASD. When we detect one or more imbalance(s), there is an increased risk that the child will receive an ASD diagnosis'
*Compared to the results of the ADOS-2 (Autism Diagnostic Observation Schedule), Second Edition
It's not clear exactly what is meant by this: it sounds as though the claim is that the blood test is more accurate than ADOS-2. That can't be right, though, because in the CAMP study, we are told: 'The Autism Diagnostic Observation Schedule–Second Version (ADOS-2) was performed by research-reliable clinicians to confirm an ASD diagnosis.' So all the ASD children in the study met ADOS-2 criteria. It looks like 'compared to' means 'based on' in this context, but it is then unclear what the 'increased risk' refers to.

How reliable is the test?
A test's validity depends crucially on its reliability: if a blood test gives different results on different occasions, then it cannot be used for diagnosis of a long-term condition. Presumably because of this, the account of the study on ClinicalTrials.gov states: 'A subset of the subjects will be asked to return to the clinic 30-60 days later to obtain a replicate metabolic profile.' Yet no data on this replicate sample is reported in the Biological Psychiatry paper.

I have no expertise in metabolomics, but it seems reasonable to suppose that amines measured in the blood may vary from one occasion to another; indeed in 2014 the authors published a preliminary report on a smaller sample from CAMP, where they specifically noted that, presumably to minimise impact of medication or special diets, blood samples were taken when the child was fasting and prior to morning administration of medication. (34% of the ASD group and 10% of the typically-developing group were on regular medication, and 19% of the ASD group were on gluten and/or casein-free diets).

I contacted the authors to ask for information on this point. They did not provide any data on test-retest reliability beyond stating:
Thirty one CAMP subjects were recruited at random for a test-retest analysis during CAMP. These subjects were all amino acid dysregulation metabotype negative at the initial time point (used in the analysis for the manuscript). The subjects were sampled 30-60 days later for retest analysis. At the second time point the 31 subjects were still metabotype negative. There are plans for additional resampling of a select group of CAMP subjects. These will include metabotype positive individuals.
Thus, we do not currently know whether a positive result on the NPDX ASD test is meaningful, in the sense of being a consistent physiological marker in the individual.

Scientific evaluation of the methods used in the Biological Psychiatry paper 
The Biological Psychiatry paper describing development of the test is highly complex, involving a wide range of statistical methods. In their previous paper with a smaller sample, the authors described thousands of blood markers and claimed that using machine learning methods, they could identify a subset that discriminated the ASD and typically-developing groups with above chance accuracy. However, they noted this finding needed confirmation in a larger sample.

In the 2018 Biological Psychiatry paper, no significant differences were found for measures of metabolite abundance, failing to replicate the 2014 findings. However, further consideration of the data led the authors to concentrate instead on ratios between metabolites. As they noted: 'Ratios can uncover biological properties not evident with individual metabolites and increase the signal when two metabolites with a negative correlation are evaluated.'

Furthermore, they focused on individuals with extreme values for ratio scores, on the grounds that ASD is a heterogeneous condition, and the interest is in identifying subgroups who may have altered metabolism. The basic logic is illustrated in Figure 1 – the idea is to find a cutoff on the distribution which selects a higher proportion of ASD than typical cases. Because 76% of the sample are ASD cases, we would expect to find 76% of cases in the tail of the distribution. However, by exploring different cutoffs, it can be possible to identify a higher proportion. The proportion of ASD cases above a positive cutoff (or below a negative cutoff) is known as the positive predictive value (PPV), and for some of the ratios examined by the researchers, it was over 90%.


Figure 1: Illustrative distributions of z-scores for 4 of the 31 metabolites in ASD and typical group: this plot shows raw levels for metabolites; blue boxes show the numbers falling above or below a cutoff that is set to maximise group differences. The final analysis focused on ratios between metabolites, rather than raw levels. From Figure S2, Smith et al (2018).

This kind of approach readily lends itself to finding spurious 'positive' results, insofar as one is first inspecting the data and then identifying a cutoff that maximises the difference between two groups. It is noteworthy that the metabolites that were selected for consideration in ratio scores were identified on the basis that they showed negative correlations within a subset of the ASD sample (the 'training set'). Accordingly, PPV values from a 'training set' are likely to be biased and will over-estimate group differences. However, to avoid circularity, one can take cutoffs from the training set, and then see how they perform with a new subset of data that was not used to derive the cutoff – the 'test set'. Provided the test set is predetermined prior to any analysis, and totally separate from the training set, then the results with the test set can be regarded as giving a good indication of how the test would perform in a new sample. This is a standard way of approaching this kind of classification problem.

Usually, the PPV for a test set will be less good than for a training set: this is just a logical consequence of the fact that observed differences between groups will involve random noise as well as true population differences, and these will boost the PPV. In the test set, random effects will be different are so are more likely to hinder rather than help prediction, and so PPV will decline. However, in the Biological Psychiatry paper, the PPVs for the test sets were only marginally different from those from the training sets: for the ratios described in Table 1, the mean PPV was .887 (range .806 - .943) for the training set, and mean .880 (range .757 - .975) for the test set.

I wanted to understand this better, and asked the authors for their analysis scripts, so I could reconstruct what they did. Here is the reply I received from Beth Donley:
We would be happy to have a call to discuss the methodology used to arrive at the findings in our paper. Our scripts and the source code they rely on are proprietary and will not be made public unless and until we publish them in a paper of our own. We think it would be more meaningful to have a call to discuss our approach so that you can ask questions and we can provide answers.
My questions were sufficiently technical and complex that this was not going to work, so I provided written questions, to which I received responses. However, although the replies were prompt, they did not really inspire confidence, and, without the scripts I could not check anything.

For instance:
My question: Is there an explanation for why the PPVs are so similar for training and test datasets? Usually you'd expect a drop in PPV in the test dataset if the function was optimised for the training dataset, just because the training threshold would inevitably be capitalising on chance.
Response: We observed this phenomenon, as well, and were surprised by the similarity of the training and test confusion matrix performance metric values. We have no way to know why the metrics were similar between sets. Our best guess is that the demographics of the training and test set of subjects had closely matched demographic and study related variables.
But the demographic similarity between a test and training set is not the main issue here. One thing that crucially determines how close the results will be is the reliability of the metabolomic measure. The lower the test-retest reliability of the measure, the more likely that results from a training set will fail to replicate. So it would be helpful if the authors would report the quantitative data that they have on this question.

If we ignore all the problems, how good is prediction? 
Unfortunately, it is virtually impossible to tell how accurate the test would be in a real-life context. First, we would have to make the assumption that a non-autistic group with developmental delay would be comparable to the typically-developing group. If non-autistic children with developmental delay show metabolomic imbalances, then the test's potential for diagnosis of ASD is compromised. Second, we would have to come up with an estimate of how many children who are given the test will actually have ASD: that's very hard to judge, but let us suppose it may be as high as 50%. Then, for the ratios reported in the Biological Psychiatry paper, we can compute that around 50% to 83% of those testing positive would have ASD. Note that the majority of children with and without ASD won't have scores in the tail of the distribution and will not therefore test positive (see Figure 1). On the NPDX website is is claimed that around 30% of children with ASD test positive: That is hard to square this with the account in Biological Psychiatry which reported 'an altered metabolic phenotype' in 16.7% of those with ASD.

Conflict of interest and need for transparency
The published paper gives a comprehensive COI statement as follows:
AMS, MAL, and REB are employees of, JJK and PRW were employees of, and ELRD is an equity owner in Stemina Biomarker Discovery Inc. AMS, JJK, PRW, MAL, ELRD, and REB are inventors on provisional patent application 62/623,153 titled “Amino Acid Analysis and Autism Subsets” filed January 29, 2018. DGA receives research funding from the National Institutes of Health, the Simons Foundation, and Stemina Biomarker Discovery Inc. He is on the scientific advisory boards of Stemina Biomarker Discovery Inc. and Axial Therapeutics.
It is generally accepted that just because there is COI, this does not invalidate the work: it simply provides a context in which it can be interpreted. The study reported in Biological Psychiatry represents a huge investment of time and money, with research funds contributed from both public and private sources. In the Xconomy interview, it is stated that the research has cost $8 million to date. This kind of work may only be possible to do with involvement of a biotechnology company which is willing to invest funds in the hope of making discoveries that can be commercialised; this is a similar model to drug development.

Where there is a strong commercial interest in the outcome of research, the best way of counteracting negative impressions is for researchers to be as open and transparent as possible. This was not the case with the NPDX study: as described above, there were substantial changes from the registered protocol on ClinicalTrials.gov, not discussed in the paper. The analysis scripts are not available – this means we have to take on trust details of the methods in an area where the devil is in the detail. As Philip Stark has argued, a paper that is long on results but short on methods is more like an advertisement than a research communication: "Science should be ‘show me’, not ‘trust me’; it should be ‘help me if you can’, not ‘catch me if you can’."

Postscript
On 27th December, Biological Psychiatry published correspondence on the Smith et al paper by Kristin Sainani and Steven Goodman from Stanford University. They raised some of the points noted above regarding the lack of predictive utility of the blood test in clinical contexts, the lack of a comparison sample with developmental delay, and the conflict of interest issues. In their response, the authors made the point that they had noted these limitations in their published paper.

References
Sainani, K. L., & Goodman, S. N. (2018). Lack of diagnostic utility of 'amino acid dysregulation metabotypes'. Biological Psychiatry. doi:10.1016/j.biopsych.2018.11.012

Smith, A. M., Donley, E. L. R., Burrier, R. E., King, J. J., & Amaral, D. G. (2018). Reply to: Lack of Diagnostic Utility of “Amino Acid Dysregulation Metabotypes”. Biological Psychiatry. doi: https://doi.org/10.1016/j.biopsych.2018.11.013

Smith, A. M., King, J., J, West, P. R., Ludwig, M. A., Donley, E. L. R., Burrier, R. E., & Amaral, D. G. (2018). Amino acid dysregulation metabotypes: Potential biomarkers for diagnosis and individualized treatment for subtypes of autism spectrum disorder. Biological Psychiatry. doi:https://doi.org/10.1016/j.biopsych.2018.08.016

Stark, P. (2018). Before reproducibiity must come preproducibility. Nature, 557, 613. doi:10.1038/d41586-018-05256-0
West, P. R., Amaral, D. G., Bais, P., Smith, A. M., Egnash, L. A., Ross, M. E., . . . Burrier, R. E. (2014). Metabolomics as a tool for discovery of biomarkers of Autism Spectrum Disorder in the blood plasma of children. PLOS One, 9(11), e112445. doi:  https://doi.org/10.1371/journal.pone.0112445

Saturday, 9 June 2018

Developmental language disorder: the need for a clinically relevant definition

There's been debate over the new terminology for Developmental Language Disorder (DLD) at a meeting (SRCLD) in the USA. I've not got any of the nuance here, but I feel I should make a quick comment on one issue I was specifically asked about, viz:


As background: the field of children's language disorders has been a terminological minefield. The term Specific Language Impairment (SLI) began to be used widely in the 1980s as a diagnosis for children who had problems acquiring language for no apparent reason. One criterion for the diagnosis was that the child's language problems should be out of line with other aspects of development, and hence 'specific', and this was interpreted as requiring normal range nonverbal IQ (nviq).

The term SLI was never adopted by the two main diagnostic systems -WHO's International Classification of Diseases (ICD) or the American Psychiatric Association's Diagnostic and Statistical Manual (DSM), but the notion that IQ should play a part in the diagnosis became prevalent.

In 2016-7 I headed up the CATALISE project with the specific goal of achieving some consensus about the diagnostic criteria and terminology for children's language disorders: the published papers about this are openly available for all to read (see below). The consensus of a group of experts from a range of professions and countries was to reject SLI in favour of the term DLD.

Any child who meets criteria for SLI will meet criteria for DLD: the main difference is that the use of an IQ cutoff is no longer part of the definition. This does not mean that all children with language difficulties are regarded as having DLD: those who meet criteria for intellectual disability, known syndromes or biomedical conditions are treated separately (see these slides for summary).

The tweet seems to suggest we should retain the term SLI, with its IQ cutoff, because it allows us to do neatly controlled research studies. I realise a brief, second-hand tweet about Rice's views may not be a fair portrayal of what she said, but it does emphasise a bone of contention that was thoroughly gnawed in the discussions of the CATALISE panel, namely, what is the purpose of diagnostic terminology? I would argue its primary purpose is clinical, and clinical considerations are not well-served by research criteria.

The traditional approach to selecting groups for research is to find 'pure' cases - quite simply, if you include children who have other problems beyond language (including other neurodevelopmental difficulties) then it is much harder to know how far you are assessing correlates or causes of language problems: things get messy and associations get hard to interpret. The importance of controlling for nonverbal IQ has been particularly emphasised over many years: quite simply, if you compare language-impaired vs comparison (typically-developing, or td) children on a language or cognitive measure, and the language-impaired group has lower nonverbal ability, then it may be that you are looking at a correlate of nonverbal ability rather than language. Restricting consideration to those who meet stringent IQ criteria to equalise the groups is one way of addressing the issue.

However, there are three big problems with this approach:

1. A child's nonverbal IQ can vary from time to time and it will depend on the test that is used. However, although this is problematic, it's not the main reason for dropping IQ cutoffs; the strongest arguments concern validity rather than reliability of an IQ-based approach.

2. The use of IQ-cutoffs ignores the fact that pure cases of language impairment are the exception rather than the rule. In CATALISE we looked at the evidence and concluded that if we were going to insist that you could only get a diagnosis of DLD if you had no developmental problems beyond language, then we'd exclude many children with language problems (see also this old blogpost). If our main purpose is to get a diagnostic system that is clinically workable, it should be applicable to the children who turn up in our clinics - not just a rarefied few who meet research criteria. An analogy can be drawn with medicine: imagine if your doctor identified you with high blood pressure but refused to treat you unless you were in every other regard fit and healthy. That would seem both unfair and ill-judged. Presence of co-occurring conditions might be important for tracking down underlying causes and determining a treatment path, but it's not a reason for excluding someone from receiving services.

3. Even for research purposes, it is not clear that a focus on highly specific disorders makes sense. An underlying assumption, which I remember starting out with, was the idea that the specific cases were in some important sense different from those who had additional problems. Yet, as noted in the CATALISE papers, the evidence for this assumption is missing: nonverbal IQ has very little bearing on a child's clinical profile, response to intervention, or aetiology. For me, what really knocked my belief in the reality of SLI as a category was doing twin studies: typically, I'd find that identical twins were very similar in their language abilities, but they sometimes differed in nonverbal ability, to the extent that one met criteria for SLI and the other did not. Researchers who treat SLI as a distinct category are at risk of doing research that has no application to the real world.

There is nothing to stop researchers focusing on 'pure' cases of language disorder to answer research questions of theoretical interest, such as questions about the modularity of language. This kind of research uses children with a language disorder as a kind of 'natural experiment' that may inform our understanding of broader issues. It is, however, important not to confuse such research with work whose goal is to discover clinically relevant information.

If practitioners let the theoretical interests of researchers dictate their diagnostic criteria, then they are doing a huge disservice to the many children who end up in a no-man's-land, without either diagnosis or access to intervention. 

References

Bishop, D. V. M. (2017). Why is it so hard to reach agreement on terminology? The case of developmental language disorder (DLD). International Journal of Language & Communication Disorders, 52(6), 671-680. doi:10.1111/1460-6984.12335

Bishop, D. V. M., Snowling, M. J., Thompson, P. A., Greenhalgh, T., & CATALISE Consortium. (2016). CATALISE: a multinational and multidisciplinary Delphi consensus study. Identifying language impairments in children. PLOS One, 11(7), e0158753. doi:10.1371/journal.pone.0158753

Bishop, D. V. M., Snowling, M. J., Thompson, P. A., Greenhalgh, T., & CATALISE Consortium. (2017). Phase 2 of CATALISE: a multinational and multidisciplinary Delphi consensus study of problems with language development: Terminology. Journal of Child Psychology and Psychiatry, 58(10), 1068-1080. doi:10.1111/jcpp.12721

Saturday, 23 August 2014

Labels for unexplained language difficulties in children: We need to talk

The view from the Tower of Babel
This week saw the publication of a special issue of the International Journal of Language and Communication Disorders, focusing on labels for children with unexplained language difficulties. Two target articles, one by Sheena Reilly and colleagues, and one by me, are accompanied by an editorial by Susan Ebbels, twenty commentaries, and a final paper where Sheena and I join forces with Bruce Tomblin to try to synthesise the different viewpoints. These articles are free for anyone to access.

Terminological battles are often boring and seldom come to any consensus, so why are we putting time into this thorny issue? Quite simply, because it really matters. As we argue in the articles, having a label affects how a children are perceived, what help they are offered, and how seriously their problems are taken. 'Specific Language Impairment' has very poor name recognition compared to dyslexia and autism, despite being at least as common. Furthermore, unless we can agree on some common language, it's difficult to make progress in research, and to discover, for instance, the underlying causes of language difficulties, how common they are in different parts of the world, or what interventions work.

I was first confronted with the full extent of the problem when I tried to analyse the amount of research and research funding associated with different developmental disorders (Bishop, 2010). There are other conditions, notably autism and dyslexia, where there is plenty of debate about diagnostic criteria, or even about whether the condition exists. But even so, the terminology is reasonably consistent. For children's language difficulties, this is not the case - they can be described as cases of language difficulty, disorder, impairment, disability, needs or delay, with various prefixes such as 'developmental', 'specific' or 'primary'. Some researchers will use such labels with precise meanings, often excluding children who have co-existing conditions, whereas others use them more descriptively. This made it extremely difficult to do a sensible internet search to estimate the amount of research funding associated with children's language difficulties.  

The confusion over labels has, I think, also contributed to the lack of public recognition of language difficulties in children. A couple of years ago, I joined together with Courtenay Norbury, Maggie Snowling, Gina Conti-Ramsden and Becky Clark with the goal of remedying this situation. We started a campaign for Raising Awareness of Language Learning Impairments (RALLI) (Bishop et al., 2012), and set up a YouTube channel to provide basic information. We spent some time debating what terminology to use: "Language learning impairment" was our preferred choice, but many of our videos talk of Specific Language Impairment, simply because that is a more familiar label. The lack of an agreed label proved a real stumbling block for our attempts at public engagement, and we decided that, as well as producing videos, one of our goals would be to get the terminology issue discussed more widely, in the hope of achieving some consensus. It was a very happy coincidence that Sheena Reilly and colleagues were crystallizing their own position on this question in an article in IJLDC, and that they, and the Editors, were willing to include my article, and the commentaries of other RALLI founders, in the published debate.

One thing that came across when reading commentaries on our articles was the disconnect between research and practice. One point on which I agree with Sheena and colleagues is that there is no justification for drawing a distinction between children whose language problems are comparable with below average nonverbal ability, and those who have a mismatch between good nonverbal skills and low language. Research has failed to find any difference between children with uneven or even nonverbal-verbal profiles in terms of responsiveness to intervention or underlying causes. Such a distinction is, however, widely used in educational and clinical settings to decide which children gain access to extra support in school.  Another issue raised by the Reilly et al paper is whether it is logical to use other exclusionary criteria, and to distinguish, for instance, between children who do and don't have autistic features in association with a language problem.  

As Susan Ebbels noted in her editorial, in everyday settings "diagnostic labels and criteria were being used creatively in disputes over access to services both by those seeking to obtain services for children (often parents and their lawyers) who could be accused of ‘diagnostic shopping’ and also by those seeking to deny services (often due to financial constraints) who may use particularly restrictive criteria in order to reduce the number of children qualifying for services". 

We can't afford to ignore this confused situation any longer. The time has come to have a wider debate on these issues, with the aim of reaching a consensus about how terms are used. The Royal College of Speech and Language Therapists has set up a moderated discussion forum where people can give their views on the best way forward. Please do consider adding your voice: it is important that all those affected by this issue have a say, whether you are a speech-language therapist/pathologist, psychologist, teacher, health professional, legal expert, policymaker, a parent of a child with language difficulties, or someone who has experienced language difficulties. We'd also love to hear from those outside the UK - whether English-speaking or not. You can access the discussion forum here.

Finally, to raise awareness of this debate, during the week of 24th-31st August I will be taking over  the @WeSpeechies Twitter handle as guest curator. On Tuesday 26th at 8.a.m. BST there will be a live twitter debate on this topic. Feel free to join in, even if you aren't a regular tweeter.

References
Bishop, D. (2010). Which Neurodevelopmental Disorders Get Researched and Why? PLoS ONE, 5 (11) DOI: 10.1371/journal.pone.0015112  
Bishop, D., Clark, B., Conti-Ramsden, G., Norbury, C., & Snowling, M. (2012). RALLI: An internet campaign for raising awareness of language learning impairments Child Language Teaching and Therapy, 28 (3), 259-262 DOI: 10.1177/0265659012459467

Slides on this topic are available here.



Addendum Friday 29th August 2014

We've had a great week of interactions on Twitter. A transcript for the week is available here.
I'll look through this and aim to organise the material in due course, but meanwhile would encourage anyone who is interested to continue the discussion on Twitter. I'm appending below some tweets that I generated throughout the week to generate debate.

As noted above, the chat links in to a special issue of the Internat. J Lang. Comm Dis which is free to access here http://t.co/ncTUaYvyoI.  NB it is not all that obvious but there are 10 commentaries after each target article.

If you want to join the discussion on Twitter, feel free to comment at any time, but, please include the #WeSpeechies hashtag, so we can aggregate comments easily. Also if your comment relates to a numbered question, please add Q1, etc so we can relate them.

Monday started with my attempt to summarise each of the  twenty commentaries in a Tweet-length message.


Summaries from commentaries

Paediatricn Gillian Baird: ICD &DSM classifications talk of 'language disorder'; implies distinct from normal variation.  Disorder’ used for conditions without obvious aetiology; functional effect described separately in ICFDH.

Lauchlan/Boyle, ed psych view. Must ask: ‘Will label change the child's life for the better? Aetiology often irrelevant

Bellair et al: community SALTs. No one label works for both research & clinical. SLI has problems but we can manage them.

Mabel Rice: "SLI has yet to receive widespread adoption in clinical practice, in spite of the great need for it." critical of DSM5: excluded "well-researched category of SLI", included SCD, "with a minimal research base"

Kate Taylor SLP. SLI underidentified. Changing the term won't resolve the issue, which is one of measurement rather than label.

Conti-Ramsden: Any Consensus Panel on terminology must be international and include voices from different languages,

Hansson et al: ICD10 labels don't map on to use by researchers in Sweden . : Sweden: phonological & grammatical difficulties seen as part of language impairment. Soc comm probs separate

Clark & Carter: Survey:Scottish SALTs unclear re terms & diagnostic criteria. Move from exclusionary to inclusionary criteria.

Hüneke & Lascelles http://t.co/9rVKJzoBZV. Concern that watering down terminology will mean kids lose scarce resources. Prefer medical term 'developmental dysphasia' that gets problems taken seriously

Strudwick/Bauer http://t.co/GSY5Xwz283 Concern that labels don't capture comorbidities; most ch with 'SLI' have other problems

Michael Rutter, psychiatrist "both clinical & research classifications needed but they require a different approach"

Rutter: Specific’ implies ‘pure’ language impairment; "not supported by any of the available evidence"

Larry Leonard: Many researchers already use broader definition of SLI: do not use term to mean children have a pure profile. communicatn with the public/other disciplines will be even harder if we adopt generic label ‘language impairment.

Snowling: DSM5 treats Communication Disorders separately from Specific Learning Disorders, yet they often co-occur

Aoife Gallagher,SALT; ethical issue:"who owns diagnosis once it has been given.. who ultimately has the right to take it away"

Andrew Whitehouse: ‘SLI’ provides neat criteria for researchers but label hides behavioural & aetiological heterogeneity

Dockrell/Lindsay Educational perspective re SLI is missing yet day-to-day support of learning/development provided by teachers. in England ‘speech, language & communication needs’ (SLCN) indicates primary need is with language & communication

Grist & Hartshorne: http://t.co/QKeQbQFsdy Children & young people we work with rarely describe selves as having SLI or SLCN

Norbury @lilaccourt Relaxing diag criteria will increase demand for services.SALTs shld focus on severe & persistent impairmts

Parsons et al @wordaware Shockwaves through SALT profession if nonverbal IQ criteria and delay/disorder distinction removed .Use of marketing approaches to development of a new term, including consultation with parents & young people.

Wright: legal perspective Much time spent in tribunal appeals arguing re labels: eg is it delay or disorder, is it specific?

Questions for debate

On Tuesday we had a live twitter chat with four question topics, and later in the week, I added further numbered question. Here is the total list – we'd love to hear your thoughts on any or all of these:

Q1 What is your view on use of the diagnostic label SLI? Does it reflect a medical model and is this appropriate.

Q2 is What are appropriate criteria for identifying children's language problems

Q3; Should IQ, ASD features, hearing loss determine whether language-impaired children can access services?

Q4 What terminology is most appropriate for children who have unexplained language problems?

Q5 ICD11 will use'Developmental Language Disorder' and DSM5 uses 'Language Disorder'. What do people think of these terms?

Q6 In research SLI still widely used but without requiring IQ discrepancy. Should we retain SLI but with this broader meaning, or is it just confusing?

Q7 In UK education, Speech, Language and Communication Needs (SLCN) is popular term. Is it used outside UK? Is it useful?

Q8 In UK clinical practice, distinction between language 'delay' & 'disorder' is used, but it has no research support.  Where does delay/disorder distinction come from? How defined?

Q9 Is there any support for a return to the more medical term 'developmental dysphasia'?

Q10. Reilly et al and several commentators suggest we drop 'Specific' and use the term 'Language Impairment' instead .What wld be advantages (e.g. avoids unfair exclusion) and disadvantages (e.g. too broad)?

Q11 What do people think of terms 'Language Learning Impairment' or 'Primary language impairment'? '

Q12 Do diagnostic labels actually help children and families?

Q13 Shld terminology/diagnostic criteria be responsibility of speechies, or shld other professions & families have a say? Assumptions/practices seem v. different in education/medicine/psychology vs speech-language therapy/pathology

Q14 In yr area, who does intervention with kids whose language problems are associated with autism?

Q15 Some  people take pride in identifying themselves as dyslexic. Does this ever happen for kids with language problems? If not, why not?

Q16 Has anyone encountered situation where child not offered intervention bcs language problems attributed to social deprivation?

Q17 Insurance considerations seldom important in UK, but affect label use elsewhere. Do US insurers just require DSM?