Showing posts with label effect size. Show all posts
Showing posts with label effect size. Show all posts

Tuesday, 5 December 2023

Low-level lasers. Part 2. Erchonia and the universal panacea

 

 


In my last blogpost, I looked at a study that claimed continuing improvements of symptoms of autism after eight 5-minute sessions where a low-level laser was pointed at the head.  The data were so extreme that I became interested in the company, Erchonia, who sponsored the study and in Regulatory Insight, Inc, whose statistician failed to notice anything odd.  In exploring Erchonia's research corpus, I found that they have investigated the use of their low-laser products for a remarkable range of conditions. A search of clinicaltrials.com with the keyword Erchonia produced 47 records, describing studies of pain (chronic back pain, post-surgical pain, and foot pain), body contouring (circumference reduction, cellulite treatment), sensorineural hearing loss, Alzheimer's disease, hair loss, acne and toenail fungus. After excluding the trials on autism described in my previous post, fourteen of the records described randomised controlled trials in which an active laser was compared with a placebo device that looked the same, with both patient and researcher being kept in the dark about which device was which until the data were analysed. As with the autism study, the research designs for these RCTs specified on clinicaltrials.com looked strong, with statistician Elvira Cawthon from Regulatory Insight involved in data analysis.

As shown in Figure 1, where results are reported for RCTs, they have been spectacular in virtually all cases. The raw data are mostly not available, and in general the plotted data look less extreme than in the autism trial covered in last week's post, but nonetheless, the pattern is a consistent one, where over half the active group meet the cutoff for improvement, whereas less than half (typically 25% or less) of the placebo group do so. 

FIGURE 1: Proportions in active treated group vs placebo group meeting preregistered criterion for improvement (Error bars show SE)*

I looked for results from mainstream science against which to benchmark the Erchonia findings.  I found a big review of behavioural and pharmaceutical interventions for obesity by the US Agency for Healthcare Research and Quality (LeBlanc et al, 2018). Figures 7 and 13 show results for binary outcomes - relative risk of losing 5% or more of body weight over a 12 month period; i.e. the proportion of treated individuals who met this criterion divided by the proportion of controls. In 38 trials of behavioural interventions, the mean RR was 1.94 [95% CI, 1.70 to 2.22]. For 31 pharmaeutical interventions, the effect varied with the specific medication, with RR ranging from 1.18 to 3.86. Only two pharmaceutical comparisons had RR in excess of 3.0. By contrast, for five trials of body contouring or cellulite reduction from Erchonia, the RRs ranged from 3.6 to 18.0.  Now, it is important to note that this is not comparing like with like: the people in the Erchonia trials were typically not clinically obese: they were mostly women seeking cosmetic improvements to their appearance.  So you could, and I am sure many would, argue it's an unfair comparison. If anyone knows of another literature that might provide a better benchmark, please let me know. The point is that the effect sizes reported by Erchonia are enormous relative to the kinds of effects typically seen with other treatments focused on weight reduction.

If we look more generally at the other results obtained with low-level lasers, we can compare them to an overview of effectiveness of common medications (Leucht et al, 2015). These authors presented results from a huge review of different therapies, with effect sizes represented as standardized mean differences (SMD - familiar to psychologists as Cohen's d). I converted Erchonia results into this metric*, and found that across all the studies of pain relief shown in Figure 1, the average SMD was 1.30, with a range from 0.87 to 1.77. This contrasts with Leucht et al's estimated effect size of 1.06 for oxycodone plus paracetamol, and 0.83 for Sumatriptan for migraine.  So if we are to believe the results, they indicate that the effect of Erchonia low-level lasers is as good or better than the most effective pharmaceutical medications that we have for pain relief or weight loss. I'm afraid I remain highly sceptical.

I would not have dreamed of looking at Erchonia's track record if it were not for their impossibly good results in the Leisman et al autism trial that I discussed in the previous blogpost.  When I looked in more detail, I was reminded of the kinds of claims made for alternative treatments for children's learning difficulties, where parents are drawn in with slick websites promising scientifically proven interventions, and glowing testimonials from satisfied customers. Back in 2012 I blogged about how to evaluate "neuroscientific" interventions for dyslexia.  Most of the points I made there apply to the world of "photomodulation" therapies, including the need to be wary when a provider claims that a single method is effective for a whole host of different conditions.  

Erchonia products are sold worldwide and seem popular with alternative health practitioners. For instance, in Stockport, Manchester, you can attend a chiropractic clinic where Zerona laser treatment will remove "stubborn body fat". In London there is a podiatry centre that reassures you: "There are numerous papers which show that cold laser affects the activity of cells and chemicals within the cell. It has been shown that cold laser can encourage the formation of stem cells which are key building blocks in tissue reparation. It also affects chemicals such as cytochrome c and causes a cascade of reactions which stimulates the healing. There is much research to show that cold laser affects healing and there are now several very good class 1 studies to show that laser can be effective." But when I looked for details of these "very good class 1 studies" they were nowhere to be found. In particular, it was hard to find research by scientists without vested interests in the technology.  

Of all the RCTs that I found, there were just two that were conducted at reputable universities. One of them, on hearing loss (NCT01820416) was conducted at the University of Iowa, but terminated prematurely because intermediate analysis showed no clinically or statistically significant effects (Goodman et al., 2013).  This contrasts sharply with NCT00787189, which had the dramatic results reported in Figure 1 (not, as far as I know, published outside of clinicaltrials.gov). The other university-based study was the autism study based in Boston described in my previous post: again, with unpublished, unimpressive results posted on clinicaltrials.gov.

This suggests it is important when evaluating novel therapies to have results from studies that are independent of those promoting the therapy. But, sadly, this is easier to recommend than to achieve. Running a trial takes a lot of time and effort: why would anyone do this if they thought it likely that the intervention would not work and the postulated mechanism of action was unproven? There would be a strong risk that you'd end up putting in effort that would end in a null result, which would be hard to publish. And you'd be unlikely to convince those who believed in the therapy - they would no doubt say you had the wrong wavelength of light, or insufficient duration of therapy, and so on.  

I suspect the response by those who believe in the power of low-level lasers will be that I am demonstrating prejudice, in my reluctance to accept the evidence that they provide of dramatic benefits. But, quite simply, if low-level laser treatment was so remarkably effective in melting fat and decreasing pain, surely it would have quickly been publicised through word of mouth from satisfied customers. Many of us are willing to subject our bodies to all kinds of punishments in a quest to be thin and/or pain-free. If this could be done simply and efficiently without the need for drugs, wouldn't this method have taken over the world?

*Summary files (Erchonia_proportions4.csv) and script (Erchonia_proportions_for_blog.R) are on Github, here.

Friday, 20 April 2012

Getting genetic effect sizes in perspective


My research focuses on neurodevelopmental disorders - specific language impairment, dyslexia, and autism in particular. For all of these there is evidence of genetic influence. But the research papers reporting relevant results are often incomprehensible to people who aren’t geneticists (and sometimes to those who are).  This leaves us ignorant of what has really been found, and subject to serious misunderstandings.
Just as preamble, evidence for genetic influences on behaviour comes in two kinds. The first approach, sometimes referred to as genetic epidemiology or behaviour genetics allows us to infer how far genes are involved in causing individual differences by studying similarities between people who have different kinds of genetic relationship. The mainstay of this field is the twin study. The logic of twin studies is pretty simple, but the methods currently used to analyse twin data are complex. The twin method is far from perfect, but it has proved useful in helping us identify which conditions are worth investigating using the second approach, molecular genetics.
Molecular genetics involves finding segments of DNA that are correlated with a behavioural (or other phenotypic) measure. It involves laboratory work analysing biological samples of people who’ve been assessed on relevant measures. So if we’re interested in, say, dyslexia, we can either look for DNA variants that predict a person’s reading ability - a quantitative approach - or we can look for DNA variants that are more common in people who have dyslexia. There’s a range of methods that can be used, depending on whether the data come from families - in which case the relationship between individuals can be taken into account - or whether we just have a sample of unrelated people who vary on the behaviour of interest, in this case reading ability.
The really big problem comes from a tendency in molecular genetics to focus just on p-values when reporting findings. This is understandable: the field of molecular genetics has been plagued by chance findings. This is because there’s vast amounts of DNA that can be analysed, and if you look at enough things, then the odd result will pop up as showing a group difference just by chance. (See this blogpost for further explanation). The p-value indicates whether an association between a DNA variant and a behavioural measure is a solid finding that is likely to replicate in another sample.
But a p-value depends on two things: (a) the strength of association between DNA and behaviour (effect size) and (b) the sample size. Psychologists, many of whom are interested in genetic variants linked to behaviour, are mostly used to working with samples that number in the tens rather than hundreds or thousands. It’s easy, therefore, to fall into the trap of assuming that a very low p-value means we have a large effect size, because that’s usually the case in the kind of studies we’re used to. Misunderstanding can arise if effect sizes are not reported in a paper.
Suppose we have a genetic locus with two alleles, a and A, and a geneticist contrasts people with an aa genotype vs those with aA or AA (who are grouped together). We read a molecular genetics paper that reports an association between these genotypes and reading ability with p-value of .001. Furthermore, we see there are other studies in the literature reporting similar associations, so this seems a robust finding. You could be forgiven for concluding that the geneticists have found a “dyslexia gene”, or at least a strong association with a DNA variant that will be useful in screening and diagnosis. And, if you are a psychologist, you might be tempted to do further studies contrasting people with aa vs aA/AA genotypes on behavioural or neurobiological measures that are relevant for reading ability.
However, this enthusiasm is likely to evaporate if you consider effect sizes. There is a nice little function in R, compute.es, that allows you to compute effect size easily if you know a p-value and a sample size. The table below shows:
  •  effect sizes (Cohen’s d, which gives mean difference between groups in z-score units)
  •  average for each group for reading scores scaled so the mean for the aA/AA group is 100 with SD of 15
Results are shown for various sample sizes with equal numbers of aa vs aA/AA and either p =.001 or p = .0001. (See reference manual for the R function for relevant formulae, which are also applicable in cases of unequal sample size). For those unfamiliar with this area, a child would not normally be flagged up as having reading problems unless a score on a test scaled this way was 85 or less (i.e., 1 SD below the mean).

Table 1: Effect sizes (Cohen’s d) and group means derived from p-value and sample size (N)










When you have the kind of sample size that experimental or clinical psychologists often work with, with 25 participants per group, a p of .001 is indicative of a big effect, with a mean difference between groups of almost one SD. However, you may be surprised at how small the effect size is when you have a large sample. If you have a sample of 3000 or so, then a difference of just 1-2 points (or .08 SD) will give you p < .001. Most molecular genetic studies have large sample sizes. Geneticists in this area have learned that they have to have large samples, because they are looking for small effects!
It would be quite wrong to suggest that only large effect sizes are interesting. Small but replicable effects can be of great importance in helping us understand causes of disorders, because if we find relevant genes we can study their mode of action (Scerri & Schulte-Korne, 2010). But, as far as the illustrative data in Table 1 are concerned, few psychologists would regard the reading deficit associated at p of .001 or .0001 with genotype aa as of clinical significance, once the sample size exceeds 1000 per group.
Genotype aa may be a risk factor for dyslexia, but only in conjunction with other risks. On its own it doesn’t cause dyslexia.  And the notion, propagated by some commercial genetics testing companies, that you could use a single DNA variant with this magnitude of effect to predict a person’s risk of dyslexia, is highly misleading.


Further reading
Flint, J., Greenspan, R. J., & Kendler, K. S. (2010). How Genes Influence Behavior: Oxford University Press.
Scerri, T., & Schulte-Körne, G. (2009). Genetics of developmental dyslexia European Child & Adolescent Psychiatry, 19 (3), 179-197 DOI: 10.1007/s00787-009-0081-0

If you are interested in analysing twin data, you can see my blog on Twin Methods in OpenMx, which illustrates a structural equation modelling approach in R with simulated data.

Update 21/4/12: Thanks to Tom Scerri for pointing out my original wording talked of "two versions of an allele", which has now been corrected to "a genetic locus with two alleles"; as Tom noted: an allele is an allele, you can't have two versions of it.
Tom also noted that in the table, I had taken the aA/AA genotype as the reference group for standardisation to mean 100, SD 15. A more realistic simulation would take the whole population with all three genotypes as the reference group, in which case the effect size would result from the aA/AA group having a mean above 100, while the aa group would have mean below 100. This would entail that, relative to the grand population average, the averages for aa would be higher than shown here, so that the number with clinically significant deficits will be even smaller.
I hope in future to illustrate these points by computing effect sizes for published molecular genetic studies reporting links with cognitive phenotypes.