Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Thursday, 12 March 2026

Bishopblog catalogue (updated 12 Mar 2026)

Source: http://www.weblogcartoons.com/2008/11/23/ideas/

Those of you who follow this blog may have noticed a lack of thematic coherence. I write about whatever is exercising my mind at the time, which can range from technical aspects of statistics to the design of bathroom taps. I decided it might be helpful to introduce a bit of order into this chaotic melange, so here is a catalogue of posts by topic.

Language impairment, dyslexia and related disorders
The common childhood disorders that have been left out in the cold (1 Dec 2010) What's in a name? (18 Dec 2010) Neuroprognosis in dyslexia (22 Dec 2010) Where commercial and clinical interests collide: Auditory processing disorder (6 Mar 2011) Auditory processing disorder (30 Mar 2011) Special educational needs: will they be met by the Green paper proposals? (9 Apr 2011) Is poor parenting really to blame for children's school problems? (3 Jun 2011) Early intervention: what's not to like? (1 Sep 2011) Lies, damned lies and spin (15 Oct 2011) A message to the world (31 Oct 2011) Vitamins, genes and language (13 Nov 2011) Neuroscientific interventions for dyslexia: red flags (24 Feb 2012) Phonics screening: sense and sensibility (3 Apr 2012) What Chomsky doesn't get about child language (3 Sept 2012) Data from the phonics screen (1 Oct 2012) Auditory processing disorder: schisms and skirmishes (27 Oct 2012) High-impact journals (Action video games and dyslexia: critique) (10 Mar 2013) Overhyped genetic findings: the case of dyslexia (16 Jun 2013) The arcuate fasciculus and word learning (11 Aug 2013) Changing children's brains (17 Aug 2013) Raising awareness of language learning impairments (26 Sep 2013) Good and bad news on the phonics screen (5 Oct 2013) What is educational neuroscience? (25 Jan 2014) Parent talk and child language (17 Feb 2014) My thoughts on the dyslexia debate (20 Mar 2014) Labels for unexplained language difficulties in children (23 Aug 2014) International reading comparisons: Is England really do so poorly? (14 Sep 2014) Our early assessments of schoolchildren are misleading and damaging (4 May 2015) Opportunity cost: a new red flag for evaluating interventions (30 Aug 2015) The STEP Physical Literacy programme: have we been here before? (2 Jul 2017) Prisons, developmental language disorder, and base rates (3 Nov 2017) Reproducibility and phonics: necessary but not sufficient (27 Nov 2017) Developmental language disorder: the need for a clinically relevant definition (9 Jun 2018) Changing terminology for children's language disorders (23 Feb 2020) Developmental Language Disorder (DLD) in relaton to DSM5 (29 Feb 2020) Why I am not engaging with the Reading Wars (30 Jan 2022)

Autism
Autism diagnosis in cultural context (16 May 2011) Are our ‘gold standard’ autism diagnostic instruments fit for purpose? (30 May 2011) How common is autism? (7 Jun 2011) Autism and hypersystematising parents (21 Jun 2011) An open letter to Baroness Susan Greenfield (4 Aug 2011) Susan Greenfield and autistic spectrum disorder: was she misrepresented? (12 Aug 2011) Psychoanalytic treatment for autism: Interviews with French analysts (23 Jan 2012) The ‘autism epidemic’ and diagnostic substitution (4 Jun 2012) How wishful thinking is damaging Peta's cause (9 June 2014) NeuroPointDX's blood test for Autism Spectrum Disorder ( 12 Jan 2019) Biomarkers to screen for autism (again) (6 Dec 2022)

Developmental disorders/paediatrics
The hidden cost of neglected tropical diseases (25 Nov 2010) The National Children's Study: a view from across the pond (25 Jun 2011) The kids are all right in daycare (14 Sep 2011) Moderate drinking in pregnancy: toxic or benign? (21 Nov 2012) Changing the landscape of psychiatric research (11 May 2014) The sinister side of French psychoanalysis revealed (15 Oct 2019) A desire for clickbait can hinder an academic journal's reputation (4 Oct 2022) Polyunsaturated fatty acids and children's cognition: p-hacking and the canonisation of false facts (4 Sep 2023)

Genetics
Where does the myth of a gene for things like intelligence come from? (9 Sep 2010) Genes for optimism, dyslexia and obesity and other mythical beasts (10 Sep 2010) The X and Y of sex differences (11 May 2011) Review of How Genes Influence Behaviour (5 Jun 2011) Getting genetic effect sizes in perspective (20 Apr 2012) Moderate drinking in pregnancy: toxic or benign? (21 Nov 2012) Genes, brains and lateralisation (22 Dec 2012) Genetic variation and neuroimaging (11 Jan 2013) Have we become slower and dumber? (15 May 2013) Overhyped genetic findings: the case of dyslexia (16 Jun 2013) Incomprehensibility of much neurogenetics research ( 1 Oct 2016) A common misunderstanding of natural selection (8 Jan 2017) Sample selection in genetic studies: impact of restricted range (23 Apr 2017) Pre-registration or replication: the need for new standards in neurogenetic studies (1 Oct 2017) Review of 'Innate' by Kevin Mitchell ( 15 Apr 2019) Why eugenics is wrong (18 Feb 2020)

Neuroscience
Neuroprognosis in dyslexia (22 Dec 2010) Brain scans show that… (11 Jun 2011)  Time for neuroimaging (and PNAS) to clean up its act (5 Mar 2012) Neuronal migration in language learning impairments (2 May 2012) Sharing of MRI datasets (6 May 2012) Genetic variation and neuroimaging (1 Jan 2013) The arcuate fasciculus and word learning (11 Aug 2013) Changing children's brains (17 Aug 2013) What is educational neuroscience? ( 25 Jan 2014) Changing the landscape of psychiatric research (11 May 2014) Incomprehensibility of much neurogenetics research ( 1 Oct 2016)

Reproducibility
Accentuate the negative (26 Oct 2011) Novelty, interest and replicability (19 Jan 2012) High-impact journals: where newsworthiness trumps methodology (10 Mar 2013) Who's afraid of open data? (15 Nov 2015) Blogging as post-publication peer review (21 Mar 2013) Research fraud: More scrutiny by administrators is not the answer (17 Jun 2013) Pressures against cumulative research (9 Jan 2014) Why does so much research go unpublished? (12 Jan 2014) Replication and reputation: Whose career matters? (29 Aug 2014) Open code: note just data and publications (6 Dec 2015) Why researchers need to understand poker ( 26 Jan 2016) Reproducibility crisis in psychology ( 5 Mar 2016) Further benefit of registered reports ( 22 Mar 2016) Would paying by results improve reproducibility? ( 7 May 2016) Serendipitous findings in psychology ( 29 May 2016) Thoughts on the Statcheck project ( 3 Sep 2016) When is a replication not a replication? (16 Dec 2016) Reproducible practices are the future for early career researchers (1 May 2017) Which neuroimaging measures are useful for individual differences research? (28 May 2017) Prospecting for kryptonite: the value of null results (17 Jun 2017) Pre-registration or replication: the need for new standards in neurogenetic studies (1 Oct 2017) Citing the research literature: the distorting lens of memory (17 Oct 2017) Reproducibility and phonics: necessary but not sufficient (27 Nov 2017) Improving reproducibility: the future is with the young (9 Feb 2018) Sowing seeds of doubt: how Gilbert et al's critique of the reproducibility project has played out (27 May 2018) Preprint publication as karaoke ( 26 Jun 2018) Standing on the shoulders of giants, or slithering around on jellyfish: Why reviews need to be systematic ( 20 Jul 2018) Matlab vs open source: costs and benefits to scientists and society ( 20 Aug 2018) Responding to the replication crisis: reflections on Metascience 2019 (15 Sep 2019) Manipulated images: hiding in plain sight (13 May 2020) Frogs or termites: gunshot or cumulative science? ( 6 Jun 2020) Open data: We know what's needed - now let's make it happen (27 Mar 2021) A proposal for data-sharing the discourages p-hacking (29 Jun 2022) Can systematic reviews help clean up science (9 Aug 2022)Polyunsaturated fatty acids and children's cognition: p-hacking and the canonisation of false facts (4 Sep 2023) Book Review: Unreliable (Csaba Szabo) (Mar 16, 2025) Gold standard science isn't gold standard if it's applied selectively - firearms (Aug 26, 2025) Gold standard science isn't gold standard if it's applied selectively - autism (Aug 27, 2025) Wellcome LEAP's new $50M program (Oct 20, 2025) The dangers of using bibliometrics with polluted data (Nov 21, 2025)  

Statistics
Book review: biography of Richard Doll (5 Jun 2010) Book review: the Invisible Gorilla (30 Jun 2010) The difference between p < .05 and a screening test (23 Jul 2010) Three ways to improve cognitive test scores without intervention (14 Aug 2010) A short nerdy post about the use of percentiles (13 Apr 2011) The joys of inventing data (5 Oct 2011) Getting genetic effect sizes in perspective (20 Apr 2012) Causal models of developmental disorders: the perils of correlational data (24 Jun 2012) Data from the phonics screen (1 Oct 2012)Moderate drinking in pregnancy: toxic or benign? (1 Nov 2012) Flaky chocolate and the New England Journal of Medicine (13 Nov 2012) Interpreting unexpected significant results (7 June 2013) Data analysis: Ten tips I wish I'd known earlier (18 Apr 2014) Data sharing: exciting but scary (26 May 2014) Percentages, quasi-statistics and bad arguments (21 July 2014) Why I still use Excel ( 1 Sep 2016) Sample selection in genetic studies: impact of restricted range (23 Apr 2017) Prospecting for kryptonite: the value of null results (17 Jun 2017) Prisons, developmental language disorder, and base rates (3 Nov 2017) How Analysis of Variance Works (20 Nov 2017) ANOVA, t-tests and regression: different ways of showing the same thing (24 Nov 2017) Using simulations to understand the importance of sample size (21 Dec 2017) Using simulations to understand p-values (26 Dec 2017) One big study or two small studies? ( 12 Jul 2018) Time to ditch relative risk in media reports (23 Jan 2020)

Journalism/science communication
Orwellian prize for scientific misrepresentation (1 Jun 2010) Journalists and the 'scientific breakthrough' (13 Jun 2010) Orwellian prize for journalistic misrepresentation: an update (29 Jan 2011) Academic publishing: why isn't psychology like physics? (26 Feb 2011) Scientific communication: the Comment option (25 May 2011)  Publishers, psychological tests and greed (30 Dec 2011) Time for academics to withdraw free labour (7 Jan 2012) 2011 Orwellian Prize for Journalistic Misrepresentation (29 Jan 2012) Communicating science in the age of the internet (13 Jul 2012) How to bury your academic writing (26 Aug 2012) Schizophrenia and child abuse in the media (26 May 2013) Why we need pre-registration (6 Jul 2013) On the need for responsible reporting of research (10 Oct 2013) Psychology research: hopeless case or pioneering field? (28 Aug 2015) When scientific communication is a one-way street (13 Dec 2016) Time to ditch relative risk in media reports (23 Jan 2020) Book Review. Fiona Fox: Beyond the Hype (12 Apr 2022)

Academic Publishing
Science journal editors: a taxonomy (28 Sep 2010) Time for neuroimaging (and PNAS) to clean up its act (5 Mar 2012) High-impact journals: where newsworthiness trumps methodology (10 Mar 2013)  A short rant about numbered journal references (5 Apr 2013) Desperate marketing from J. Neuroscience ( 18 Feb 2016) Editorial integrity: publishers on the front line ( 11 Jun 2016) A New Year's letter to academic publishers (4 Jan 2014) Journals without editors: What is going on? (1 Feb 2015) Editors behaving badly? (24 Feb 2015) Will Elsevier say sorry? (21 Mar 2015) How long does a scientific paper need to be? (20 Apr 2015) Will traditional science journals disappear? (17 May 2015) My collapse of confidence in Frontiers journals (7 Jun 2015) Publishing replication failures (11 Jul 2015) Breaking the ice with buxom grapefruits: Pratiques de publication and predatory publishing (25 Jul 2017) Should editors edit reviewers? ( 26 Aug 2018) Corrigendum: a word you may hope never to encounter (3 Aug 2019) Percent by most prolific author score and editorial bias (12 Jul 2020) PEPIOPs – prolific editors who publish in their own publications (16 Aug 2020) Faux peer-reviewed journals: a threat to research integrity (6 Dec 2020) Time for publishers to consider the rights of readers as well as authors (13 Mar 2021) Universities vs Elsevier: who has the upper hand? (14 Nov 2021) We need to talk about editors (6 Sep 2022) So do we need editors? (11 Sep 2022) Reviewer-finding algorithms: the dangers for peer review (30 Sep 2022) A desire for clickbait can hinder an academic journal's reputation (4 Oct 2022) What is going on in Hindawi special issues? (12 Oct 2022) New Year's Eve Quiz: Dodgy journals special (31 Dec 2022) A suggestion for e-Life (20 Mar 2023) Papers affected by misconduct: Erratum, correction or retraction? (11 Apr 2023) Is Hindawi “well-positioned for revitalization?” (23 Jul 2023) The discussion section: Kill it or reform it? (14 Aug 2023) Spitting out the AI Gobbledegook sandwich: a suggestion for publishers (2 Oct 2023) The world of Poor Things at MDPI journals (Feb 9 2024) Some thoughts on eLife's New Model: One year on (Mar 27 2024) Does Elsevier's negligence pose a risk to public health? (Jun 20 2024) Collapse of scientific standards at MDPI journals: a case study (Jul 23 2024) My experience as a reviewer for MDPI (Aug 8 2024) Optimizing research integrity investigations: the need for evidence (Aug 22 2024) Now you see it, now you don't: the strange world of disappearing Special Issues at MDPI (Sep 4 2024) Prodding the behemoth with a stick (Sep 14 2024) Using PubPeer to screen editors (Sep 24 2024) An open letter regarding Scientific Reports (Oct 16 2024) What's going on at the Journal of Psycholinguistic Research? (Oct 21, 2024) Finland vs Germany: the case of MDPI (Dec 23, 2024) Tomatoes roaming the fields: another embarrassing paper for MDPI (Jan 18, 2025) IEEE has a pseudoscience problem (Feb 22, 2025) Trouble at t'(review) mill: How MDPI lets down authors (July 21 ,2025) New publishing models will only work if authors embrace them (July 31, 2025) Problems with eLife's new article type: Replication studies (Oct 27, 2025) The inner workings of a paper mill (Nov 8, 2025) An open letter to the BMJ editorial board (Jan 5, 2026) An analysis of PubPeer comments on highly-cited retracted articles (Feb 2, 2026) Stealth corrections are still a threat to academic integrity (Feb 20, 2026) 

Social Media
A gentle introduction to Twitter for the apprehensive academic (14 Jun 2011) Your Twitter Profile: The Importance of Not Being Earnest (19 Nov 2011) Will I still be tweeting in 2013? (2 Jan 2012) Blogging in the service of science (10 Mar 2012) Blogging as post-publication peer review (21 Mar 2013) The impact of blogging on reputation ( 27 Dec 2013) WeSpeechies: A meeting point on Twitter (12 Apr 2014) Email overload ( 12 Apr 2016) How to survive on Twitter - a simple rule to reduce stress (13 May 2018)

Academic life
An exciting day in the life of a scientist (24 Jun 2010) How our current reward structures have distorted and damaged science (6 Aug 2010) The challenge for science: speech by Colin Blakemore (14 Oct 2010) When ethics regulations have unethical consequences (14 Dec 2010) A day working from home (23 Dec 2010) Should we ration research grant applications? (8 Jan 2011) The one hour lecture (11 Mar 2011) The expansion of research regulators (20 Mar 2011) Should we ever fight lies with lies? (19 Jun 2011) How to survive in psychological research (13 Jul 2011) So you want to be a research assistant? (25 Aug 2011) NHS research ethics procedures: a modern-day Circumlocution Office (18 Dec 2011) The REF: a monster that sucks time and money from academic institutions (20 Mar 2012) The ultimate email auto-response (12 Apr 2012) Well, this should be easy…. (21 May 2012) Journal impact factors and REF2014 (19 Jan 2013)  An alternative to REF2014 (26 Jan 2013) Postgraduate education: time for a rethink (9 Feb 2013)  Ten things that can sink a grant proposal (19 Mar 2013)Blogging as post-publication peer review (21 Mar 2013) The academic backlog (9 May 2013)  Discussion meeting vs conference: in praise of slower science (21 Jun 2013) Why we need pre-registration (6 Jul 2013) Evaluate, evaluate, evaluate (12 Sep 2013) High time to revise the PhD thesis format (9 Oct 2013) The Matthew effect and REF2014 (15 Oct 2013) The University as big business: the case of King's College London (18 June 2014) Should vice-chancellors earn more than the prime minister? (12 July 2014)  Some thoughts on use of metrics in university research assessment (12 Oct 2014) Tuition fees must be high on the agenda before the next election (22 Oct 2014) Blaming universities for our nation's woes (24 Oct 2014) Staff satisfaction is as important as student satisfaction (13 Nov 2014) Metricophobia among academics (28 Nov 2014) Why evaluating scientists by grant income is stupid (8 Dec 2014) Dividing up the pie in relation to REF2014 (18 Dec 2014)  Shaky foundations of the TEF (7 Dec 2015) A lamentable performance by Jo Johnson (12 Dec 2015) More misrepresentation in the Green Paper (17 Dec 2015) The Green Paper’s level playing field risks becoming a morass (24 Dec 2015) NSS and teaching excellence: wrong measure, wrongly analysed (4 Jan 2016) Lack of clarity of purpose in REF and TEF ( 2 Mar 2016) Who wants the TEF? ( 24 May 2016) Cost benefit analysis of the TEF ( 17 Jul 2016)  Alternative providers and alternative medicine ( 6 Aug 2016) We know what's best for you: politicians vs. experts (17 Feb 2017) Advice for early career researchers re job applications: Work 'in preparation' (5 Mar 2017) Should research funding be allocated at random? (7 Apr 2018) Power, responsibility and role models in academia (3 May 2018) My response to the EPA's 'Strengthening Transparency in Regulatory Science' (9 May 2018) More haste less speed in calls for grant proposals ( 11 Aug 2018) Has the Society for Neuroscience lost its way? ( 24 Oct 2018) The Paper-in-a-Day Approach ( 9 Feb 2019) Benchmarking in the TEF: Something doesn't add up ( 3 Mar 2019) The Do It Yourself conference ( 26 May 2019) A call for funders to ban institutions that use grant capture targets (20 Jul 2019) Research funders need to embrace slow science (1 Jan 2020) Should I stay or should I go: When debate with opponents should be avoided (12 Jan 2020) Stemming the flood of illegal external examiners (9 Feb 2020) What can scientists do in an emergency shutdown? (11 Mar 2020) Stepping back a level: Stress management for academics in the pandemic (2 May 2020)
TEF in the time of pandemic (27 Jul 2020) University staff cuts under the cover of a pandemic: the cases of Liverpool and Leicester (3 Mar 2021) Some quick thoughts on academic boycotts of Russia (6 Mar 2022) When there are no consequences for misconduct (16 Dec 2022) Open letter to CNRS (30 Mar 2023) When privacy rules protect fraudsters (Oct 12, 2023) Defence against the dark arts: a proposal for a new MSc course (Nov 19, 2023) An (intellectually?) enriching opportunity for affiliation (Feb 2 2024) Just make it stop! When will we say that further research isn't needed? (Mar 24 2024) Are commitments to open data policies worth the paper they are written on? (May 26 2024) Whistleblowing, research misconduct, and mental health (Jul 1 2024) I don't care about journal impact factors but I do care about visibility (Oct 27, 2024) Why I have resigned from the Royal Society (Nov 25, 2024) Seven reasons for keeping Elon Musk as a Fellow of the Royal Society (Feb 12, 2025) The dangers of using bibliometrics with polluted data (Nov 21, 2025)

Celebrity scientists/quackery
Three ways to improve cognitive test scores without intervention (14 Aug 2010) What does it take to become a Fellow of the RSM? (24 Jul 2011) An open letter to Baroness Susan Greenfield (4 Aug 2011) Susan Greenfield and autistic spectrum disorder: was she misrepresented? (12 Aug 2011) How to become a celebrity scientific expert (12 Sep 2011) The kids are all right in daycare (14 Sep 2011)  The weird world of US ethics regulation (25 Nov 2011) Pioneering treatment or quackery? How to decide (4 Dec 2011) Psychoanalytic treatment for autism: Interviews with French analysts (23 Jan 2012) Neuroscientific interventions for dyslexia: red flags (24 Feb 2012) Why most scientists don't take Susan Greenfield seriously (26 Sept 2014) NeuroPointDX's blood test for Autism Spectrum Disorder ( 12 Jan 2019) Low-level lasers. Part 1. Shining a light on an unconventional treatment for autism (Nov 25, 2023) Low-level lasers. Part 2. Erchonia and the universal panacea (Dec 5, 2023)

Women
Academic mobbing in cyberspace (30 May 2010) What works for women: some useful links (12 Jan 2011) The burqua ban: what's a liberal response (21 Apr 2011) C'mon sisters! Speak out! (28 Mar 2012) Psychology: where are all the men? (5 Nov 2012) Should Rennard be reinstated? (1 June 2014) How the media spun the Tim Hunt story (24 Jun 2015)

Politics and Religion
Lies, damned lies and spin (15 Oct 2011) A letter to Nick Clegg from an ex liberal democrat (11 Mar 2012) BBC's 'extensive coverage' of the NHS bill (9 Apr 2012) Schoolgirls' health put at risk by Catholic view on vaccination (30 Jun 2012) A letter to Boris Johnson (30 Nov 2013) How the government spins a crisis (floods) (1 Jan 2014) The alt-right guide to fielding conference questions (18 Feb 2017) We know what's best for you: politicians vs. experts (17 Feb 2017) Barely a good word for Donald Trump in Houses of Parliament (23 Feb 2017) Do you really want another referendum? Be careful what you wish for (12 Jan 2018) My response to the EPA's 'Strengthening Transparency in Regulatory Science' (9 May 2018) What is driving Theresa May? ( 27 Mar 2019) A day out at 10 Downing St (10 Aug 2019) Voting in the EU referendum: Ignorance, deceit and folly ( 8 Sep 2019) Harry Potter and the Beast of Brexit (20 Oct 2019) Attempting to communicate with the BBC (8 May 2020) Boris bingo: strategies for (not) answering questions (29 May 2020) Linking responsibility for climate refugees to emissions (23 Nov 2021) Response to Philip Ball's critique of scientific advisors (16 Jan 2022) Boris Johnson leads the world ....in the number of false facts he can squeeze into a session of PMQs (20 Jan 2022) Some quick thoughts on academic boycotts of Russia (6 Mar 2022) Contagion of the political system (3 Apr 2022)When there are no consequences for misconduct (16 Dec 2022)

Humour and miscellaneous Orwellian prize for scientific misrepresentation (1 Jun 2010) An exciting day in the life of a scientist (24 Jun 2010) Science journal editors: a taxonomy (28 Sep 2010) Parasites, pangolins and peer review (26 Nov 2010) A day working from home (23 Dec 2010) The one hour lecture (11 Mar 2011) The expansion of research regulators (20 Mar 2011) Scientific communication: the Comment option (25 May 2011) How to survive in psychological research (13 Jul 2011) Your Twitter Profile: The Importance of Not Being Earnest (19 Nov 2011) 2011 Orwellian Prize for Journalistic Misrepresentation (29 Jan 2012) The ultimate email auto-response (12 Apr 2012) Well, this should be easy…. (21 May 2012) The bewildering bathroom challenge (19 Jul 2012) Are Starbucks hiding their profits on the planet Vulcan? (15 Nov 2012) Forget the Tower of Hanoi (11 Apr 2013) How do you communicate with a communications company? ( 30 Mar 2014) Noah: A film review from 32,000 ft (28 July 2014) The rationalist spa (11 Sep 2015) Talking about tax: weasel words ( 19 Apr 2016) Controversial statues: remove or revise? (22 Dec 2016) The alt-right guide to fielding conference questions (18 Feb 2017) My most popular posts of 2016 (2 Jan 2017) An index of neighbourhood advantage from English postcode data ( 15 Sep 2018) Working memories: A brief review of Alan Baddeley's memoir ( 13 Oct 2018) New Year's Eve Quiz: Dodgy journals special (31 Dec 2022) Retrospective look at blog highlights of 2024 (Jan 1, 2025)

Monday, 27 July 2020

TEF in the time of pandemic

An article in the Times Higher today considers the fate of the Teaching Excellence Framework (TEF).  I am a long-term critic of the TEF, on the grounds that it lacks an adequate rationale,  has little statistical or content validity, is not cost-effective, and has the potential to mislead potential students about the quality of teaching in higher education institutions. For a slideshow covering these points, see here. I was pleased to be quoted in the Times Higher article, alongside other senior figures in higher education, who were in broad agreement that the future of TEF now seems uncertain. Here I briefly document three of my concerns.

First, the fact that the Pearce Review has not been published is reminiscent of the Government's strategy of sitting on reports that it finds inconvenient. I think we can assume the report is not a bland endorsement of TEF, but rather that it did identify some of the fundamental statistical problems with the methodology of TEF, all of which just get worse when extended down to subject-level TEF. My own view is that subject-level TEF would be unworkable. If this is what the report says, then it would be an embarrassment for government, and a disappointment for universities who have already invested in the exercise. I'm not confident that this would stop TEF going ahead, but this may be a case where after so many changes of minister, the government would be willing to either shelve the idea (the more sensible move) or just delay in the hope they can overcome the problems.

Second, the whole nature of teaching has changed radically in response to the pandemic. Of course, we are all uncertain of the future, and institutions vary in terms of their predictions, but what I am hearing from the experts in pandemics is that it is wrong to imagine we are living through a blip after which we will return to normal. Some staff are adapting well to the demand for online teaching, but this is going to depend on how far teaching requires a practical element, as well as on how tech-savvy individual teaching staff are. So, if much teaching stays online, then we'd be evaluating universities on a very different teaching profile than the one assessed in TEF.

Finally, there is wide variation in how universities are responding to the impact of the pandemic on staff. Some are making staff redundant, especially those on short-term contracts, and many are in financial difficulties. Jobs are being frozen. Even in well-established universities such as my own, there are significant numbers of staff who are massively impacted by having children to care for at home. Overall, what it means is that the teaching that is delivered is not only different in kind, but actual and effective staff/student ratios are likely to go down.

So my bottom line is that even if the TEF methodology worked (and it doesn't), it's not clear that the statistics used for it would be relevant in future. I get the impression that some HEIs are taking the approach that the show must go on, with regard to both REF and TEF, because they have substantial sunk costs in these exercises (though more for REF than TEF). But staff are incredibly hard-pressed in just delivering teaching and I think enthusiasm for TEF, never high, is at rock bottom right now. 

At the annual lecture of the Council for Defence of British Universities in 2018 I argued that TEF should have been strangled at birth. It has struggled on in a sickly and miserable state since 2015. It is now time to put it out of its misery.

Sunday, 12 January 2020

Should I stay or should I go? When debate with opponents should be avoided

Suppose you are invited to speak at a conference where some of the other speakers have views very different from yours. What do you do? My guess is that most academics would say you should accept. After all, we progress by evaluating claims and counterclaims, and robust debate is the lifeblood of scientific research. I'm going to argue here that there are exceptions and explain why I think responsible scientists should avoid a meeting called "Fixing Science: Practical Solutions for the Irreproducibility Crisis".

To understand this reaction, it helps to have read Merchants of Doubt: How a Handful of Scientists Obscured the Truth on Issues from Tobacco Smoke to Global Warming by Eric Conway and Naomi Oreskes (reviewed here). The synopsis from the book blurb is as follows:
The U.S. scientific community has long led the world in research on such areas as public health, environmental science, and issues affecting quality of life. Our scientists have produced landmark studies on the dangers of DDT, tobacco smoke, acid rain, and global warming. But at the same time, a small yet potent subset of this community leads the world in vehement denial of these dangers.

Merchants of Doubt tells the story of how a loose-knit group of high-level scientists and scientific advisers, with deep connections in politics and industry, ran effective campaigns to mislead the public and deny well-established scientific knowledge over four decades. Remarkably, the same individuals surface repeatedly - some of the same figures who have claimed that the science of global warming is "not settled" denied the truth of studies linking smoking to lung cancer, coal smoke to acid rain, and CFCs to the ozone hole. "Doubt is our product," wrote one tobacco executive. These 'experts' supplied it.
Uncertainty about science that threatens big businesses has been promoted by think tanks such as the Heartland Institute and Cato Institute, which receive substantial funding from those vested interests. The Fixing Science meeting has a clear overlap with those players.

The meeting first came to my attention when a mini Twitterstorm erupted after James Heathers tweeted:
Everyone's familiar with the manel, right? The all-male panel? Well, here's a whole new level for you.
Presenting: THE MANFERENCE.

I found myself wondering whether the lack of women was deliberate – maybe the organisers think you need a Y chromosome to do science – or whether it just showed they were tone-deaf to current social norms.

Once I twigged that this was organised by the National Association of Scholars, then everything fell into place. You can find the background to this organisation in their annual report here.

As several commentators have pointed out, they use the acronym NAS, which just happens to be the same as the highly respectable National Academy of Sciences: to avoid confusion here I will refer to them as NatAsSchols. The impression from their website and publications is that they are aligned with a neoliberal viewpoint and are opposed to attempts to increase diversity of race or gender in Universities.

So why is this organisation, whose mission is focused on issues such as preserving free speech and counteracting left-wing bias in Universities, running a meeting on fixing problems in science?

The NatAsSchols explains their interest in this topic as follows:
In April we published The Irreproducibility Crisis, a report on the modern scientific crisis of reproducibility—the failure of a shocking amount of scientific research to discover true results because of slipshod use of statistics, groupthink, and flawed research techniques. We launched the report at the Rayburn House Office Building in Washington, DC; it was introduced by Representative Lamar Smith, the Chairman of the House Committee on Science, Space, and Technology. This project signals our increasing commitment to address the academy’s flawed science as well as its abandonment of Western civilization and the liberal arts. We are following up The Irreproducibility Crisis with the investigation of four government agencies, including the Environmental Protection Agency. We are determined to find out just how badly irreproducible science has distorted government policy.
This makes it clear that the agenda is fundamentally a political one, designed to support the Trump administration's dismantling of environmental protections.

In 2018, Naomi Oreskes, author of Merchants of Doubt, wrote in Nature about a new 'Transparency Rule' proposed by the Environmental Protection Agency:
There is a crisis in US science, but it is not the one claimed by advocates for the rule. The crisis is the attempt to discredit scientific findings that threaten powerful corporate interests. The EPA is following a pattern that I and others have documented in regard to tobacco smoke, pollution, climate, and more. One tactic exploits the idea of scientific uncertainty to imply there is no scientific consensus. Another, seen in the latest efforts, insinuates that relevant research might be flawed. To add insult to injury, those using these tactics claim to be defending science.
February's meeting is in the same mould. The format of the meeting is cleverly constructed. The conference will be introduced and summed up by David J. Theroux (Founder, President and Chief Executive Officer of the Independent Institute and Publisher of The Independent Review) and Peter Wood, (President, NatAsSchols). Neither man has any scientific background. Theroux delighted the Heartland Institute last summer when he promoted the idea, recently publicised by Donald Trump, that wind turbines are responsible for killing numerous birds (to see this lampooned, click here)

Wood was an anthropologist who has been Provost at a small religious school, The King’s College in New York City (2005-2007), before moving to NatAsSchols. He has, as far as I can tell, no peer-reviewed publications, but he has written pieces deriding climate concerns, e.g. "the fantasies of global warming catastrophe are a kind of substitute religion, replete with a salvation doctrine, rituals of expiation, and a collection of demons to be cast out."

Another presenter is David Randall, who is Director of Research at NatAsSchols, policy advisor to the Heartland Institute and first author of the report on "The Irreproducibility of Modern Science". He is an unusual person to be authoring an authoritative report on the state of science. Web of Science turned up seven publications by him, all in politics journals, and none with any citations. His background is in history, library studies and fiction writing.

A rather puzzling choice of speaker is Richard K. Vedder, Distinguished Emeritus Professor of Economics, Ohio University and senior fellow at The Independent Institute, a think-tank founded by David Theroux. I could not find much evidence that he has shown any prior interest in science. He is Founding Director of the Center for College Affordability and Productivity in Washington, D.C and policy advisor to the Heartland Institute.

But there are also some accredited scientists on the programme, who can be divided into two camps. First, we have a set of five speakers who are aligned with NatAsSchols and/or the Heartland Institute and who have unconventional views on subjects such as climate change, pollution and gay relationships:

Elliott D. Bloom is Professor Emeritus at the Kavli Institute for Particle Astrophysics and Cosmology in the Stanford Linear Accelerator Laboratory (SLAC) and a Fellow of the American Physical Society. He has an entry on the Independent Institute website which states: "He was a member of the SLAC team with Jerome I. Friedman, Henry W. Kendall and Richard E. Taylor who received the 1990 Nobel Prize in Physics." I thought this meant he was a Nobel Laureate, but he's not listed as one. Nevertheless, he has a strong publication record in Physics. He has co-authored a presentation on "Global Warming: Fact or Fiction?", which concludes that the sun, rather than CO2 is the principal driver of climate change.

Anastasios Tsonis is Emeritus Distinguished Professor, Department of Mathematical Sciences, Atmospheric Sciences Group, University of Wisconsin Milwaukee; and Adjunct Research Scientist, Hydrologic Research Center, San Diego, California. He has worked on mathematical models of atmospheric processes and has a strong set of publications. He is a member of the academic advisory council of the Global Warming Policy Forum, a think tank founded by Nigel Lawson to combat policies designed to mitigate climate change.

Patrick J. Michaels, Senior Fellow, Competitive Enterprise Institute has a Wikipedia entry that states that "he was a senior fellow in environmental studies at the Cato Institute until Spring 2019. Until 2007 he was research professor of environmental sciences at the University of Virginia, where he had worked from 1980." Michaels also has an entry in the Website of the Heartland Institute 

Louis Anthony Cox is Professor, Department of Biostatistics and Informatics, University of Colorado and President of Cox Associates, a Denver-based applied research company specializing in quantitative risk analysis, causal modeling, advanced analytics, and operations research. He has a long list of publications. A Google search turns up an article in the Los Angeles Times which states:
The Trump administration’s reliance on industry-funded environmental specialists is again coming under fire, this time by researchers who say that Louis Anthony 'Tony' Cox Jr., who leads a key Environmental Protection Agency advisory board on air pollution, is a 'fringe' scientist and ideologue pushing policies detrimental to public health.
They refer to this paper in Science, which stated that Cox ignored consensus viewpoints on the effects of smog and particulate pollution. His work has also been criticised for its conflict with corporate interests.

Mark Regnerus, Professor, Sociology Department, University of Texas at Austin has a Wikipedia page which notes the controversy around his research on the adverse impact of a child having a parent who has been involved in a same-sex relationship. The research is funded by the Witherspoon Institute, a conservative think tank. Regnerus also contributed to an amicus brief in opposition to same-sex marriage. A sympathetic account of the controversy was published by the NatAsSchols .

Of the remaining 11 speakers, as far as I can see only one, Barry Smith (University at Buffalo), has any formal links with NatAsSchols. With a few exceptions they are psychologists/philosophers/statisticians with a specific interest in scientific reproducibility. Interest in this topic has been growing exponentially over the past 10 years, and, in general, those engaged in research in this area do so with the aim of improving scientific transparency and practice. However, they run the risk that their agenda can be weaponised to cast doubt on any particular part of scientific research that is politically or commercially inconvenient.

They will serve perfectly as foils to the five speakers whose minority views on climate/pollution/sexuality will not have to face questioning by anyone with deep expertise in those areas. I have no doubt that the reproducibility experts will have lively debate among themselves as to how the irreproducibility crisis should be fixed, in the process achieving the useful (to the organisers) goal of emphasising just what an unreliable and uncertain business science is.

Should they agree to speak at this meeting? As @briandavidearp, remarked on Twitter: "I'm wary of deliberately failing to engage/interact w/ people or organizations on the basis that they have diff moral or political commitments than me. That way balkanization and polarization lies."

That's an answer I would have agreed with a few years ago – after all, isn't that what academic life is all about? We should not just sit in our own bubble; rather we should engage respectfully with those who have different views. But this is really not about regular scientific debate. It's about weaponising the reproducibility debate to bolster the message that everything in science is uncertain – which is very convenient for those who wish to promote fringe ideas.

My view is that many of the speakers at this meeting are being played. On the one hand their presence on the programme may encourage other to agree to participate, and give false reassurance to attendees that this is a regular conference. And on the other, they will find that their arguments are scooped up by the Merchants of Doubt and used to argue that science is so uncertain that we should not accept the consensus view. We cannot be sure whether anthropogenic climate change is exaggerated, whether pollution is not really harmful, and whether gay relationships are damaging. Those who are concerned to see such ideas promoted without any debate between experts in those areas may wish to reconsider whether this meeting is really about 'Fixing Science', or whether it is rather about 'Fitting up Scientists'.

P.S.
15th January 2020
I thank Lee Jussim for engaging in the comments below. I can see that, given his understanding of the situation, it would make sense to take part in the meeting. But his understanding is different from mine, and I want to add this PS to clarify what I'm saying.

It was perhaps a mistake for me to note the neoliberal affiliations of NatAsSchols, as this appears to have given Lee the impression that my objections to the Fixing Science meeting is based on disapproval of talking to those with right-wing beliefs. I am myself on the left politically, but I agree with Brian Earp that, insofar as it is possible to do so in good faith, we should engage with those with differing views. If the meeting consisted solely of experts in philosophy of sciences/methods/metascience, with different political persuasions, I would not be warning people off – quite the contrary.

Indeed, the people I identified as belonging to the second group of speakers would seem to be exactly such a group. As I noted, I doubt that they will come to a consensus about how to 'Fix Science', but a good mix of views and perspectives is represented. No doubt Lee and others will discuss an issue that is of particular interest to me, which is how our social and cognitive biases affect how we evaluate evidence (see Bishop, 2020). Such biases are not specifically associated with left- or right-leaning politics – they affect all of us.

The problem I have with the meeting is not that the organisers are right-wing, but rather that their organisation's goals are linked to issues around higher education, and they have no credentials in science, yet they fervently advocate minority views about such topics as climate change.  Consider how bizarre it would be if, for instance, the Psychonomics Society declared that it planned to hold a meeting on 'Fixing Politics'. The NatAsScholars just doesn't have credibility in the area of scientific pratices. Alas, what they do have instead are links with funders whose vast wealth is used to attack science that threatens their vested interests. In this respect, I think the argument that 'the left-wingers are just as bad' breaks down.

But, I reiterate, the main point is not whether NatAsSchlols is left- or right-wing. It's the weird structuring of the meeting, which juxtaposes a set of experts in the 'reproducibility crisis' with a set of individuals who promote scientific views that are far from mainstream. The fact that the topics are ones that are supported by the Heartland Institute is telling, but the same strategy could in principle be used with any fringe view. Suppose you were a sceptic about evolution or vaccination, or a believer in pre-cognition. You know your arguments would not survive scrutiny by experts familiar with evidence in the area, so you don't invite those (and to be fair, it's unlikely that they'd come anyway, as there are diminishing returns in engaging with those whose minds are fixed). But what you can do is to cast doubt on all scientific evidence by inviting along those who are questioning the solidity and credibility of current scientific practices. That's what is happening here.

The general strategy has been in use for years, as documented by Conway and Oreskes, and applied to diverse topics such as tobacco dangers and acid rain, as well as climate change. The Merchants of Doubt love it when scientists themselves disagree about the nature of evidence, because it gives them a get-out-of-jail-free card.

I'm firmly of the belief we should not shove problems with science under the carpet: we need to understand the nature and extent of such problems in order to fix them. But it is a mistake to engage with those who want to exploit the presence of uncertainty to give credibility to their fringe views.

Bishop, D. V. M. (2020). The psychology of experimental psychologists: Overcoming cognitive constraints to improve research. The 45th Sir Frederic Bartlett Lecture Quarterly Journal of Experimental Psychology, 73(1), 1-19. doi:10.1177/1747021819886519

Sunday, 3 March 2019

Benchmarking in the TEF: Something doesn't add up (v.2)





Update: March 6th 
This is version 2 of this blogpost, taking into account new insights into the weird z-scores used in TEF.  I had originally suggested there might be an algebraic error in the formula used to derive z-scores: I now realise there is a simpler explanation, which is that the z-scores used in TEF are not calculated in the usual way, with the standard deviation as denominator, but rather with the standard error of measurement as denominator. 
In exploring this issue, I've greatly benefited from working openly with a R markdown script on Github, as that has allowed others with statistical expertise to propose alternative analyses and explanations. This process is continuing, and those interested in technical details can follow developments as they happen on Github, see benchmarking_Feb2019.rmd.
Maybe my experience will encourage OfS to adopt reproducible working practices.


I'm a long-term critic of the Teaching Excellence and Student Outcomes Framework (TEF). I've put forward a swathe of arguments against the rationale for TEF in this lecture, as well as blogging for the Council for Defence of British Universities (CDBU) about problems with its rationale and statistical methods. But this week, things got even more interesting. In poking about in the data behind the TEF, I stumbled upon some anomalies that suggest to me that the TEF is not just misguided, but also is based on a foundation of statistical error.

Statistical critiques of TEF are not new. This week, the Royal Statistical Society wrote a scathing report on the statistical limitations of TEF, complaining that their previous evidence to TEF evaluations had been ignored, and stating: 'We are extremely worried about the entire benchmarking concept and implementation. It is at the heart of TEF and has an inordinately large influence on the final TEF outcome'. They expressed particular concern about the lack of clarity regarding the benchmarking methodology, which made it impossible to check results.

This reflects concerns I have had, which have led me to do further analyses of the publicly available TEF datasets. The conclusion I have come to is that the way in which z-scores are defined is very different from the usual interpretation, and leads to massive overdiagnosis of under- and over-performing institutions.

Needless, to say, this is all quite technical, but even if you don't follow the maths, I suggest you just consider the analyses reported below, in which I compare the benchmarking output from the Draper and Gittoes method with that from an alternative approach.

Draper & Gittoes (2004): a toy example

Benchmarking is intended to provide a way of comparing institutions on some metric, while taking into account differences between institutions in characteristics that might be expected to affect their performance, such as the subjects of study, and the social backgrounds of students. I will refer to these as 'contextual factors'.

The method used to do benchmarking comes from Draper and Gittoes, 2004, and is explained in this document by the Higher Education Statistics Agency: HESA. A further discussion of the method can be found in this pdf of slides from a talk by Draper (2006).

Draper (2006) provides a 'small world' example with 5 universities and 2 binary contextual categories, age and gender, to yield four combinations of contextual factors. The numbers in the top part of the chart are the proportions in each contextual (PCF) category meeting the criterion of student continuation.  The numbers in the bottom part are the numbers of students in each contextual category.


Table 1. Small world example from Draper 2006, showing % passing benchmark (top) and N students (bottom)

Essentially, the obtained score (weighted mean column) for an institution is an average of indicator values for each combination of contextual factors, weighted by the numbers with each combination of contextual factors in the institution. The benchmarked score is computed by taking the average score for each combination across all institutions (bottom row of top table) and then for each institution creating a mean score, weighted by the number in each category for a that institution. Though cumbersome (and hard to explain in words!) it is not difficult to compute.  You can find an R markdown script that does the computation here (see benchmarking_Feb2019.rmd, benchmark_function). The difference between obtained values and benchmarked value can then be computed, to see if the institution is scoring above expectation (positive difference) or below expectation (negative difference).  Results for the small world example are shown in Table 2.
Table 2. Benchmarks (Ei) computed for small world example
The column headed Oi is the observed proportion with a pass mark on the indicator (student continuation), Ei is the benchmark (expected) value for each institution, and Di is the difference between the two.

Computing standard errors of difference scores

The next step is far more complex. A z-score is computed by dividing the difference between observed and expected values on an indicator (Di) by a denominator, which is variously referred to as a standard deviation and a standard error in the documents on benchmarking.

For those who are not trained in statistics, the basic logic here is that the estimate of an institution's performance will be more labile if it is based on a small sample. If the institution takes on only 5 students each year, then estimates of completion rates from year to year will be variable - in a year where one student drops out, then the completion rate is only 80%, but if none drop out it will be 100%. You would not expect it to be constant because of random factors outside the control of the institution will affect student drop-outs. In contrast, for an institution with 1000 students, we will see much less variation from year to year. The standard error provides an estimate of the extent to which we expect the estimate of average drop-out to vary from year to year, taking size of population into account. 

To interpret benchmarked scores we need a way of estimating the standard error of the difference between the observed score on a metric (such as completion rate) and the benchmarked score, reflecting how much we would expect this to vary from one occasion to another. Only then can we judge whether the institution's performance is in line with expectation. 

Draper (2006) walks the reader through a standard method for computing the standard errors, based on the rather daunting formulae of Figure 1. The values in the SE column of table 2 are computed this way, and the z-scores are obtained by dividing each Di value by its corresponding SE.

Fomulae 5 to 8 are used to compute difference scores and standard errors (Draper, 2006)
Now anyone familiar with z-scores will notice something pretty odd about the values in Table 2. The absolute z-scores given by this method seem remarkably large: In this toy example, we see z-scores with absolute values of 5, 9 and 20.  Usually z-scores range from about -3 to 3. (Draper noted this point).


Z-scores in real TEF data

Next, I downloaded some real TEF data, so I could see whether the distribution of z-scores was unusual. Data from Year 2 (2017) in .csv format can be downloaded from this website.
The z-scores here have been computed by HESA. Here is the distribution of core z-scores for one of the metrics (Non-continuation) for the 233 institutions with data on FullTime students.

The distribution is completely out of line with what we would expect from a z-score distribution.  Absolute z-scores greater than 3, which should be vanishingly rare, are common - with the exact number varying across the six available metrics, but ranging from 33% to 58%.

Yet, they are interpreted in TEF as if a large z-score is an indicator of abnormally good or poor performance:

From p. 42  of this pdf giving Technical Specifications:

"In TEF metrics the number of standard deviations that the indicator is from the benchmark is given as the Z-score. Differences from a benchmark with a Z-score +/-1.9623 will be considered statistically significant. This is equivalent to a 95% confidence interval (that is, we can have 95% confidence that the difference is not due to chance)."

What does the z-score represent?

Z-scores feature heavily in my line of work: in psychological assessment they are used to identify people whose problems are outside the normal range. However, they aren't computed like the TEF z-scores, because they involve dividing a mean score by the standard deviation, rather than by the standard error.


It's easiest to explain this by an analogy. I'm 169 cm tall. Suppose you want to find out if that's out of line with the population of women in Oxford. You measure 10,000 women and find their mean height is 170 cm, with a standard deviation of 3. On a conventional z-score, my height is unremarkable. You just divide the difference between my height and the population height and divide by the standard deviation, -1/3, to give a z-score of -0.33. That's well within the normal limits used by TEF of -1.96 to 1.96.

Now let's compute the standard error of the population mean - to do that we compute the standard error, which is the standard deviation divided by the square root of the sample size, which gives 3/100 or .03. From that information we can get an estimate of the precision of our estimate of the population mean: we multiply the SE by 1.96, and add and subtract that value to the mean to get 95% confidence limits, which are 169.94 and 170.06. If we were to compute the z-score corresponding to my height using the SE instead of the SD, I would seem to be alarmingly short: the value would be -1/.03 = -33.33.

So what does that mean? Well the second z-score based on the SE does not test whether my height is in line with the population of 10,000 women. It tests whether my height can be regarded as equivalent to that of the average from that population. Because the population is very large, the estimate of the average is very precise, and my height is outside the error of measurement for the mean.

The problem with the TEF data is that they use the latter, SE-based method to evaluate differences from the benchmark value, but appear to interpret it as if it was a conventional SD-based z-score:

E.g. in the Technical Specificiations document (5.63):

As a test of the likelihood that a difference between a provider’s benchmark and its indicator is due to chance alone, a z-score +/- 3.0 means the likelihood of the difference being due to chance alone has reduced substantially and is negligible.

As illustrated with the height analogy, the SE-based method seems designed to over-identify high and low-achieving institutions. The only step taken to counteract this trend is an ad hoc one: because large institutions are particularly prone to obtain extreme scores, a large absolute z-score is only flagged as 'significant' if the absolute difference score is also greater than 2 or 3 percentage points. Nevertheless, the number of flagged institutions for each metric, is still far higher than would be the case if conventional z-scores based on the SD were used.

Relationship between SE-based and SD-based z-scores
(N.B. Thanks to Jonathan Mellon who noted an error in my script for computing the true z-scores. 
This update and correction made 20.20 p.m. on 6 March 2019).

I computed conventional z-scores by dividing each institution's difference from benchmark by the SD of for difference scores for all institutions and plotted it against the TEF z-scores. An example for one of the metrics is shown below. The range is in line with expectation (most values between -3 and +3) for the conventional z-scores, but much bigger for the TEF z-scores.




Conversion of z-scores into flags

In TEF benchmarking, TEF z-scores are converted into 'flags', ranging from - - or -, to denote performance below expectation, up to + or ++ for performance above expectation, with = used to indicate performance in line with expectation. It is these flags that the TEF panel considers when deciding which award (Gold, Silver or Bronze) to award.

Draper-Gittoes z-scores are flagged for significance as follows:
  •  - - z-score of -3 or less, AND an absolute difference between observed and expected values of 3%. 
  •  - z-score of -2 or less, AND an absolute difference between observed and expected values of 2%. 
  •  + z-score of 2 or more, AND an absolute difference between observed and expected values of 2%.   
  • ++ z-score of 3 or more, AND an absolute difference between observed and expected values of 3%. 
Given the problems with the method outlined above, this method is likely to massively overdiagnose both problems and good performance.

Using quantiles rather than TEF z-scores

Given that the z-scores obtained with the Draper-Gittoes method are so extreme, it could be argued that flags should be based on quantiles rather than z-score cutoffs, omitting the additional absolute difference criterion. For instance, for the Year 2 TEF data (Core z-scores) we can find cutoffs corresponding to the most extreme 5% or 1%.  If flags were based on these, then we would award extreme flags (- - or ++) only to those with negative z-scores of -13.7 or less, or positive score of 14.6 or more; less extreme flags would be awarded to those with negative z-score of -7 or less (- flag), or positive z-score of 8.6 or more (+).

Update 6th March: An alternative way of achieving the same end would be to use the TEF cutoffs with conventional z-scores; this would achieve a very similar result.

Deriving award levels from flags

It is interesting to consider how this change in procedure would affect the allocation of awards. In TEF, the mapping from raw data to awards is complex and involves more than just a consideration of flags: qualitative information is also taken into account. Furthermore, as well as the core metrics, which we have looked at above, the split metrics are also considered - i.e. flags are also awarded for subcategories, such as male/female, disabled/non-disabled: in all there are around 130 flags awarded across the six metrics for each institution. But not all flags are treated equally: the three metrics based on the National Student Survey are given half the weight of other metrics.

Not surprisingly, if we were to recompute flag scores based on quantiles, rather than using the computed z-scores, the proportions of institutions with Bronze or Gold awards drops massively.

When TEF awards were first announced, there was a great deal of publicity around the award of Bronze to certain high-profile institutions, in particularly the London School of Economics, Southampton University, University of Liverpool and the School of Oriental and African Studies. On the basis of quantile scores for Core metrics, none of these would meet criteria for Bronze: their flag scores would be -1, 0, -.5 and 0 respectively. But these are not the only institutions to see a change in award when quantiles are used. The majority of smaller institutions awarded Bronze obtain flag scores of zero.

The same is true of Gold Awards. Most institutions that were deemed to significantly outperform their benchmarks no longer do so if quantiles are used.

Conclusion

Should we therefore change the criteria used in benchmarking and adopt quantile scores? Because I think there are other conceptual problems with benchmarking, and indeed with TEF in general, I would not make that recommendation. I would prefer to see TEF abandoned. I hope the current analysis can at least draw people's attention to the questionable use of statistics used in deriving z-scores and their corresponding flags. The difference between a Bronze, Silver and Gold can potentially have a large impact on an institution's reputation. The current system for allocating these awards is not, to my mind, defensible.

I will, of course, be receptive to attempts to defend it or to show errors in my analysis, which is fully documented with scripts on github, benchmarking_Feb2019.rmd.