Showing posts with label TEF. Show all posts
Showing posts with label TEF. Show all posts

Sunday, 3 March 2019

Benchmarking in the TEF: Something doesn't add up (v.2)





Update: March 6th 
This is version 2 of this blogpost, taking into account new insights into the weird z-scores used in TEF.  I had originally suggested there might be an algebraic error in the formula used to derive z-scores: I now realise there is a simpler explanation, which is that the z-scores used in TEF are not calculated in the usual way, with the standard deviation as denominator, but rather with the standard error of measurement as denominator. 
In exploring this issue, I've greatly benefited from working openly with a R markdown script on Github, as that has allowed others with statistical expertise to propose alternative analyses and explanations. This process is continuing, and those interested in technical details can follow developments as they happen on Github, see benchmarking_Feb2019.rmd.
Maybe my experience will encourage OfS to adopt reproducible working practices.


I'm a long-term critic of the Teaching Excellence and Student Outcomes Framework (TEF). I've put forward a swathe of arguments against the rationale for TEF in this lecture, as well as blogging for the Council for Defence of British Universities (CDBU) about problems with its rationale and statistical methods. But this week, things got even more interesting. In poking about in the data behind the TEF, I stumbled upon some anomalies that suggest to me that the TEF is not just misguided, but also is based on a foundation of statistical error.

Statistical critiques of TEF are not new. This week, the Royal Statistical Society wrote a scathing report on the statistical limitations of TEF, complaining that their previous evidence to TEF evaluations had been ignored, and stating: 'We are extremely worried about the entire benchmarking concept and implementation. It is at the heart of TEF and has an inordinately large influence on the final TEF outcome'. They expressed particular concern about the lack of clarity regarding the benchmarking methodology, which made it impossible to check results.

This reflects concerns I have had, which have led me to do further analyses of the publicly available TEF datasets. The conclusion I have come to is that the way in which z-scores are defined is very different from the usual interpretation, and leads to massive overdiagnosis of under- and over-performing institutions.

Needless, to say, this is all quite technical, but even if you don't follow the maths, I suggest you just consider the analyses reported below, in which I compare the benchmarking output from the Draper and Gittoes method with that from an alternative approach.

Draper & Gittoes (2004): a toy example

Benchmarking is intended to provide a way of comparing institutions on some metric, while taking into account differences between institutions in characteristics that might be expected to affect their performance, such as the subjects of study, and the social backgrounds of students. I will refer to these as 'contextual factors'.

The method used to do benchmarking comes from Draper and Gittoes, 2004, and is explained in this document by the Higher Education Statistics Agency: HESA. A further discussion of the method can be found in this pdf of slides from a talk by Draper (2006).

Draper (2006) provides a 'small world' example with 5 universities and 2 binary contextual categories, age and gender, to yield four combinations of contextual factors. The numbers in the top part of the chart are the proportions in each contextual (PCF) category meeting the criterion of student continuation.  The numbers in the bottom part are the numbers of students in each contextual category.


Table 1. Small world example from Draper 2006, showing % passing benchmark (top) and N students (bottom)

Essentially, the obtained score (weighted mean column) for an institution is an average of indicator values for each combination of contextual factors, weighted by the numbers with each combination of contextual factors in the institution. The benchmarked score is computed by taking the average score for each combination across all institutions (bottom row of top table) and then for each institution creating a mean score, weighted by the number in each category for a that institution. Though cumbersome (and hard to explain in words!) it is not difficult to compute.  You can find an R markdown script that does the computation here (see benchmarking_Feb2019.rmd, benchmark_function). The difference between obtained values and benchmarked value can then be computed, to see if the institution is scoring above expectation (positive difference) or below expectation (negative difference).  Results for the small world example are shown in Table 2.
Table 2. Benchmarks (Ei) computed for small world example
The column headed Oi is the observed proportion with a pass mark on the indicator (student continuation), Ei is the benchmark (expected) value for each institution, and Di is the difference between the two.

Computing standard errors of difference scores

The next step is far more complex. A z-score is computed by dividing the difference between observed and expected values on an indicator (Di) by a denominator, which is variously referred to as a standard deviation and a standard error in the documents on benchmarking.

For those who are not trained in statistics, the basic logic here is that the estimate of an institution's performance will be more labile if it is based on a small sample. If the institution takes on only 5 students each year, then estimates of completion rates from year to year will be variable - in a year where one student drops out, then the completion rate is only 80%, but if none drop out it will be 100%. You would not expect it to be constant because of random factors outside the control of the institution will affect student drop-outs. In contrast, for an institution with 1000 students, we will see much less variation from year to year. The standard error provides an estimate of the extent to which we expect the estimate of average drop-out to vary from year to year, taking size of population into account. 

To interpret benchmarked scores we need a way of estimating the standard error of the difference between the observed score on a metric (such as completion rate) and the benchmarked score, reflecting how much we would expect this to vary from one occasion to another. Only then can we judge whether the institution's performance is in line with expectation. 

Draper (2006) walks the reader through a standard method for computing the standard errors, based on the rather daunting formulae of Figure 1. The values in the SE column of table 2 are computed this way, and the z-scores are obtained by dividing each Di value by its corresponding SE.

Fomulae 5 to 8 are used to compute difference scores and standard errors (Draper, 2006)
Now anyone familiar with z-scores will notice something pretty odd about the values in Table 2. The absolute z-scores given by this method seem remarkably large: In this toy example, we see z-scores with absolute values of 5, 9 and 20.  Usually z-scores range from about -3 to 3. (Draper noted this point).


Z-scores in real TEF data

Next, I downloaded some real TEF data, so I could see whether the distribution of z-scores was unusual. Data from Year 2 (2017) in .csv format can be downloaded from this website.
The z-scores here have been computed by HESA. Here is the distribution of core z-scores for one of the metrics (Non-continuation) for the 233 institutions with data on FullTime students.

The distribution is completely out of line with what we would expect from a z-score distribution.  Absolute z-scores greater than 3, which should be vanishingly rare, are common - with the exact number varying across the six available metrics, but ranging from 33% to 58%.

Yet, they are interpreted in TEF as if a large z-score is an indicator of abnormally good or poor performance:

From p. 42  of this pdf giving Technical Specifications:

"In TEF metrics the number of standard deviations that the indicator is from the benchmark is given as the Z-score. Differences from a benchmark with a Z-score +/-1.9623 will be considered statistically significant. This is equivalent to a 95% confidence interval (that is, we can have 95% confidence that the difference is not due to chance)."

What does the z-score represent?

Z-scores feature heavily in my line of work: in psychological assessment they are used to identify people whose problems are outside the normal range. However, they aren't computed like the TEF z-scores, because they involve dividing a mean score by the standard deviation, rather than by the standard error.


It's easiest to explain this by an analogy. I'm 169 cm tall. Suppose you want to find out if that's out of line with the population of women in Oxford. You measure 10,000 women and find their mean height is 170 cm, with a standard deviation of 3. On a conventional z-score, my height is unremarkable. You just divide the difference between my height and the population height and divide by the standard deviation, -1/3, to give a z-score of -0.33. That's well within the normal limits used by TEF of -1.96 to 1.96.

Now let's compute the standard error of the population mean - to do that we compute the standard error, which is the standard deviation divided by the square root of the sample size, which gives 3/100 or .03. From that information we can get an estimate of the precision of our estimate of the population mean: we multiply the SE by 1.96, and add and subtract that value to the mean to get 95% confidence limits, which are 169.94 and 170.06. If we were to compute the z-score corresponding to my height using the SE instead of the SD, I would seem to be alarmingly short: the value would be -1/.03 = -33.33.

So what does that mean? Well the second z-score based on the SE does not test whether my height is in line with the population of 10,000 women. It tests whether my height can be regarded as equivalent to that of the average from that population. Because the population is very large, the estimate of the average is very precise, and my height is outside the error of measurement for the mean.

The problem with the TEF data is that they use the latter, SE-based method to evaluate differences from the benchmark value, but appear to interpret it as if it was a conventional SD-based z-score:

E.g. in the Technical Specificiations document (5.63):

As a test of the likelihood that a difference between a provider’s benchmark and its indicator is due to chance alone, a z-score +/- 3.0 means the likelihood of the difference being due to chance alone has reduced substantially and is negligible.

As illustrated with the height analogy, the SE-based method seems designed to over-identify high and low-achieving institutions. The only step taken to counteract this trend is an ad hoc one: because large institutions are particularly prone to obtain extreme scores, a large absolute z-score is only flagged as 'significant' if the absolute difference score is also greater than 2 or 3 percentage points. Nevertheless, the number of flagged institutions for each metric, is still far higher than would be the case if conventional z-scores based on the SD were used.

Relationship between SE-based and SD-based z-scores
(N.B. Thanks to Jonathan Mellon who noted an error in my script for computing the true z-scores. 
This update and correction made 20.20 p.m. on 6 March 2019).

I computed conventional z-scores by dividing each institution's difference from benchmark by the SD of for difference scores for all institutions and plotted it against the TEF z-scores. An example for one of the metrics is shown below. The range is in line with expectation (most values between -3 and +3) for the conventional z-scores, but much bigger for the TEF z-scores.




Conversion of z-scores into flags

In TEF benchmarking, TEF z-scores are converted into 'flags', ranging from - - or -, to denote performance below expectation, up to + or ++ for performance above expectation, with = used to indicate performance in line with expectation. It is these flags that the TEF panel considers when deciding which award (Gold, Silver or Bronze) to award.

Draper-Gittoes z-scores are flagged for significance as follows:
  •  - - z-score of -3 or less, AND an absolute difference between observed and expected values of 3%. 
  •  - z-score of -2 or less, AND an absolute difference between observed and expected values of 2%. 
  •  + z-score of 2 or more, AND an absolute difference between observed and expected values of 2%.   
  • ++ z-score of 3 or more, AND an absolute difference between observed and expected values of 3%. 
Given the problems with the method outlined above, this method is likely to massively overdiagnose both problems and good performance.

Using quantiles rather than TEF z-scores

Given that the z-scores obtained with the Draper-Gittoes method are so extreme, it could be argued that flags should be based on quantiles rather than z-score cutoffs, omitting the additional absolute difference criterion. For instance, for the Year 2 TEF data (Core z-scores) we can find cutoffs corresponding to the most extreme 5% or 1%.  If flags were based on these, then we would award extreme flags (- - or ++) only to those with negative z-scores of -13.7 or less, or positive score of 14.6 or more; less extreme flags would be awarded to those with negative z-score of -7 or less (- flag), or positive z-score of 8.6 or more (+).

Update 6th March: An alternative way of achieving the same end would be to use the TEF cutoffs with conventional z-scores; this would achieve a very similar result.

Deriving award levels from flags

It is interesting to consider how this change in procedure would affect the allocation of awards. In TEF, the mapping from raw data to awards is complex and involves more than just a consideration of flags: qualitative information is also taken into account. Furthermore, as well as the core metrics, which we have looked at above, the split metrics are also considered - i.e. flags are also awarded for subcategories, such as male/female, disabled/non-disabled: in all there are around 130 flags awarded across the six metrics for each institution. But not all flags are treated equally: the three metrics based on the National Student Survey are given half the weight of other metrics.

Not surprisingly, if we were to recompute flag scores based on quantiles, rather than using the computed z-scores, the proportions of institutions with Bronze or Gold awards drops massively.

When TEF awards were first announced, there was a great deal of publicity around the award of Bronze to certain high-profile institutions, in particularly the London School of Economics, Southampton University, University of Liverpool and the School of Oriental and African Studies. On the basis of quantile scores for Core metrics, none of these would meet criteria for Bronze: their flag scores would be -1, 0, -.5 and 0 respectively. But these are not the only institutions to see a change in award when quantiles are used. The majority of smaller institutions awarded Bronze obtain flag scores of zero.

The same is true of Gold Awards. Most institutions that were deemed to significantly outperform their benchmarks no longer do so if quantiles are used.

Conclusion

Should we therefore change the criteria used in benchmarking and adopt quantile scores? Because I think there are other conceptual problems with benchmarking, and indeed with TEF in general, I would not make that recommendation. I would prefer to see TEF abandoned. I hope the current analysis can at least draw people's attention to the questionable use of statistics used in deriving z-scores and their corresponding flags. The difference between a Bronze, Silver and Gold can potentially have a large impact on an institution's reputation. The current system for allocating these awards is not, to my mind, defensible.

I will, of course, be receptive to attempts to defend it or to show errors in my analysis, which is fully documented with scripts on github, benchmarking_Feb2019.rmd.


Friday, 17 February 2017

We know what's best for you: politicians vs. experts

-->
I regard politicians as a much-maligned group. The job is not, after all, particularly well paid, when you consider the hours that they usually put in, the level of scrutiny they are subjected to, and the high-stakes issues they must grapple with. I therefore start with the assumption that most of them go into politics because they feel strongly about social or economic issues and want to make a difference. Although being a politician gives you some status, it also inevitably means you will be subjected to abuse or worse. The murder of Jo Cox led to a brief lull in the hostilities, but it's resumed with a vengeance as politicians continue to grapple with issues that divide the nation and that people feel strongly about. It seems inevitable, then, that anyone who stays the course must have the hide of a rhinoceros, and so by a process of self-selection, politicians are a relatively tough-minded lot. 

I fear, though, that in recent years, as the divisions between parties have become more extreme, so have the characteristics of politicians. One can admire someone who sticks to their principles in the face of hostile criticism; but what we now have are politicians who are stubborn to the point of pig-headedness, and simply won't listen to evidence or rational argument. So loath are they to appear wavering, that they dismiss the views of experts.

This was most famously demonstrated by the previous justice secretary, Michael Gove, who, when asked if any economists backed Brexit, replied "people in this country have had enough of experts". This position is continued by Theresa May as she goes forth in the quest for a Hard Brexit.

Then we have the case of the Secretary of State for Health, Jeremy Hunt, who has repeatedly ignored expert opinion on the changes he has introduced to produce a 'seven-day NHS'. The evidence he cited for the need for the change was misrepresented, according to the authors of the report, who were unhappy with how their study was being used. The specific plans Hunt proposed were described as 'unfunded, undefined and wholly unrealistic' by the British Medical Association, yet he pressed on.

At a time when the NHS is facing staff shortages, and as Brexit threatens to reduce the number of hospital staff from the EU, he has introduced measures that have led to demoralisation of junior doctors. This week he unveiled a new rota system that has a mix of day and night shifts that had doctors, including experts in sleep, up in arms. It was suggested that this kind of rota would not be allowed in the aviation industry, and is likely to put the health of doctors as well as patients at risk.
A third example comes from academia, where Jo Johnson, Minister of State for Universities, Science, Research and Innovation, steadfastly refuses to listen to any criticisms of his Higher Education and Research Bill, either from academics or from the House of Lords. Just as with Hunt and the NHS, he starts from fallacious premises – the idea that teaching is often poor, and that students and employers are dissatisfied – and then proceeds to introduce measures that are designed to fix the apparent problem, but which are more likely to damage a Higher Education system which, as he notes, is currently the envy of the world. The use of the National Student Survey as a metric for teaching excellence has come under particularly sharp attack – not just because of poor validity, but also because the distribution of scores make it unsuited for creating any kind of league table: a point that has been stressed by the Royal Statistical Society, the Office for National Statistics, and most recently by Lord Lipsey, joint chair of the All Party Statistics Group.

Johnson's unwillingness to engage with the criticism was discussed recently at the Annual General Meeting of the Council for Defence of British Universities (where Martin Wolf gave a dazzling critique of the Higher Education and Research Bill from an expert economics perspective).  Lord Melvyn Bragg said that in years of attending the House of Lords he had never come across such resistance to advice. I asked whether anyone could explain why Johnson was so obdurate. After all, he is presumably a highly intelligent man, educated at one of our top Universities. It's clear that he is ideologically committed to a market in higher education, but presumably he doesn't want to see the UK's international reputation downgraded, so why doesn't he listen to the kind of criticism put forward in the official response to his plans by Cambridge University? I don't know the answer, but there are two possible reasons that seem plausible to me.

First, those who are in politics seldom seem to understand the daily life of people affected by the Bills they introduce. One senior academic told me that Oxford and Cambridge in particular do themselves a disservice when they invite senior politicians to an annual luxurious college feast, in the hope of gaining some influence. The guest may enjoy the exquisite food and wine, but they go away convinced that all academics are living the high life, and give only the occasional lecture between bouts of indulgence. Any complaints, thus, are seen as those coming from idle dilettantes who are out of touch with the real world and alarmed at the idea they may be required to do serious work. Needless to say, this may have been accurate in the days of Brideshead Revisited, but it could not be further from the truth today – in Higher Education Institutions of every stripe, academics work longer hours than the average worker (though fewer, it must be said, than the hard-pressed doctors).

Second, governments always want to push things through because if they don't, they miss a window of opportunity during their period in power. So there can be a sense of, let's get this up and running and worry about the detail later. That was pretty much the case made by David Willetts when the Bill was debated in the House of Lords:

These are not perfect measures. We are on a journey, and I look forward to these metrics being revised and replaced by superior metrics in the future. They are not as bad as we have heard in some of the caricatures of them, and in my experience, if we wait until we have a perfect indicator and then start using it, we will have a very long wait. If we use the indicators that we have, however imperfect, people then work hard to improve them. That is the spirit with which we should approach the TEF today.

However, that is little comfort to those who might see their University go out of business while the problems are fixed. As Baroness Royall said in response:

My Lords, the noble Lord, Lord Willetts, said that we are embarking on a journey, which indeed we are, but I feel that the car in which we will travel does not yet have all the component parts. I therefore wonder if, when we have concluded all our debates, rather than going full speed ahead into a TEF for everybody who wants to participate, we should have some pilots. In that way the metrics could be amended quite properly before everybody else embarks on the journey with us.

Much has been said about the 'post-truth' age in which we now live, where fake news flourishes and anyone's opinion is as good as anyone else's. If ever there was a need for strong universities as a source of reliable, expert evidence, it is now. Unless academics start to speak out to defend what we have, it is at risk of disappearing.

For more detail of the case against the TEF, see here.

Saturday, 6 August 2016

Alternative providers and alternative medicine

Jo Johnson thinks that the market in Higher Education is unfair. There are currently stringent procedures in place for any institution that wants to award degrees and call itself a University. One facet of this is the requirement that any new provider must initially have its degrees validated by an established University; The University of Suffolk, which gained University title this month, provides an example of how this can work, but many of those seeking to enter the market are unhappy with current arrangements.  Roxanne Stockwell, Principal of Pearson College, complained: “Until an institution has its own degree-awarding powers, it cannot offer degrees without being validated by an existing university. Under the current system, a new partner has to find a willing validating partner, and it is locked out if it cannot.”

Jo Johnson, in his role as Minister for Universities and Science last year, criticised this arrangement. He memorably said: “I know some validation relationships work well, but the requirement for new providers to seek out a suitable validating body from amongst the pool of incumbents is quite frankly anti-competitive. It’s akin to Byron Burger having to ask permission of McDonald’s to open up a new restaurant. .....I can announce that we will shortly be lifting the moratorium that has been in place for applications for new Degree Awarding Powers and for University Title. Once again, we are opening the doors to new entrants and challenger institutions, all in the interest of increasing the choices available to students.”

So how wide should the door be opened? This raises deep questions about what constitutes a University. Currently, British Universities have a strong international reputation. Historically, they have been subject to strict scrutiny in return for receiving government funding. They combine research and teaching to push forward the boundaries of knowledge, and have trained students to value knowledge for its own sake, not just as a means to an end.

Not everyone wants that kind of education: some students do not enjoy formal academic study and may prefer vocational courses or apprenticeships. It is important that our higher education system caters for them. The question confronting us now is how far we should extend the definition of a University. It is interesting to consider the Word Cloud I made of ‘alternative providers’, shown in Figure 1.

Figure 1: Alternative providers from http://www.hefce.ac.uk/reg/register/getthedata/. Key: Blue have University Title; Brown have Degree-Awarding Powers; Pink offer designated courses; Violet deliver HE as a franchise only. Taken from http://cdbu.org.uk/speedy-entrances-and-sharp-exits-letting-in-more-alternative-providers/.
Among those listed are six institutions providing training in various forms of complementary and alternative medicine. Two of these had progressed to the point of having degree-awarding powers, which is one step below having University Title.

Let us be clear: the subject matter that these institutions teach is not endorsed by serious scientists.  David Colquhoun highlighted a worrying trend for UK Universities to give degrees in ‘anti-science’ back in 2007. In Australia, where chiropractic has been taught within the regular University system, it has come under increasing attack by scientists, who note that by awarding degrees in these subjects, one gives credibility to procedures that are placebos at best and dangerous at worst.

Are these concerns just a sign of anti-competitiveness and elitism? Are scientists trying to squeeze out the alternative providers because they think they’ll poach students from them, just as McDonalds might try to ensure that Byron Burgers are denied space for development? I’d argue not, and furthermore that Jo Johnson’s analogy shows a startling lack of understanding of what a University is all about. Medicine has had many false leads and it would be the height of arrogance to assume that what we know now is the only truth. But the difference between medicine and alternative medicine is not just that there is evidence for effectiveness for medicine; it’s also that in medicine there is a continuous movement to develop and improve theory and practice, rigorously testing and debating ideas and using scientific methods to evaluate them. It’s difficult to do this well and it often goes wrong, but there is broad agreement about the importance of evidence.

Well, you might say, what about religion, another topic that features heavily among alternative providers? Should theology be banned from Universities because it is not evidence-based? The answer is no, and for similar reasons to those given above: the difference between our traditional Universities and the new providers is that Universities teaching theology consider a range of perspectives and teach students critical thinking. You do not need to be a believer in any God to study theology. In contrast, new providers are often narrow in their focus: many of those with religious affiliations look as if they train students in one religious viewpoint and one only.

Defining the difference between what is suitable and not suitable for inclusion in a University degree course is itself an interesting intellectual exercise. We should not assume that something is good just because it already exists, and is bad if it is new.  But if we just ignore this distinction and have a free-for-all whereby anything can be regarded as higher education provided that there are students willing to pay for it, we will end up with a system in which the terms University and Degree will count for very little, and where the survival of a higher education institution has more to do with its marketing skills than its academic standing. In the past, there were few institutions clamouring to become Universities because it was not easy to make a profit from Higher Education. That has all changed now that higher education providers can get their hands on money from the Student Loans Company. Experience to date suggests we need to have more, rather than less, scrutiny of alternative providers in the current financial climate.

One final point: the alternative providers I have discussed here could do very well on the Teaching Excellence Framework (TEF), which will rate higher education institutions according to three main criteria: student satisfaction, drop-out rates and employability. Indeed, if they recruit their students from among disadvantaged social groups, they might well achieve higher TEF scores than more selective institutions, because benchmarking is used to adjust the outcomes. So under Jo Johnson’s oversight, we could end up with a situation where the quality of teaching at the Anglo-European College of Chiropractic is deemed superior to that at the Universities of Oxford and Cambridge. An interesting thought.

Sunday, 17 July 2016

Cost-benefit analysis of the Teaching Excellence Framework

©CartoonStock.com
The government’s new Higher Education and Research Bill gets its second reading this week. One complaint is that it has been rushed in without adequate scrutiny of some key components. I was interested, therefore, to discover, that a Detailed Impact Assessment was published in June, specifically to look at the costs and benefits of the various components of the Bill. What I found was quite shocking: we were being told that the financial benefits of the new Teaching Excellence Framework (TEF) vastly outweighed its costs – yet look in detail and this is all smoke and mirrors.

In particular, the report shows that while the costs of TEF to the higher education sector (confusingly described as ‘business’) are estimated at £20 million, the direct benefits will come to £1,146 million, giving a net benefit of £1,126 million (Table 1). How could the introduction of a new bureaucratic evaluation exercise be so remarkably beneficial? I read on with bated breath.

Well, sad to relate, it’s voodoo analysis.  This becomes clear if you press on to Table 12, which shows the crucial data from statistical modelling. Quite simply, the TEF generates money for institutions that get a good rating because it allows them to increase fees in line with inflation. Institutions that don’t participate in the TEF, or those that fail to get a good enough rating, will not be able to exceed the current 9K per annum fee, and so in real terms their income will decline over time. As far as I can make out, they are not included in Table 1. Furthermore, the increases for the compliant, successful institutions are measured relative to how they would have done if they had not been allowed to raise fees.

So to sum up:
  • You don’t need the TEF to achieve this result. You could get the same outcome by just allowing all institutions to raise fees in line with inflation.
  • As noted in the briefing to the Bill by the House of Commons: “the Bill is expected to result in a net financial benefit to higher education providers of around £1.1billion a year. This is in very large part due to the higher fees that providers with successful TEF outcomes will be able to charge students.” (p. 59)
  • The system is designed for there to be winners and losers, and the losers will inevitably see their real income falling further and further behind the winners, unless inflation is zero.
The impact assessment does consider other options, including that of allowing fee increases in line with inflation provided the institution has a satisfactory Quality Assurance rating. This is rejected on the grounds that: “whilst QA is a good starting point, reliance on QA alone and in the longer-term will not enable significant differentiation of teaching quality to help inform student decisions and encourage institutions to improve their teaching quality.” (p. 37).  This makes clear that one consequence (and one suspects one purpose) of TEF is to facilitate the division into institutional sheep and goats, followed by starvation of the goats.

Another option, which was strongly recommended by many of those who responded to the consultation exercise on the Green Paper which preceded the bill, is to remove the link between TEF and fees. In other words, have some kind of teaching evaluation, where the motivation for taking part would be reputational rather than financial.  This too is rejected as not sufficiently powerful an incentive: “the Research Excellence Framework allocates £1.5bn a year to institutions. To achieve parity of esteem and focus between teaching and research the TEF will need to have a similar level of financial implications.” However, this is rather disingenuous. There is no pot of money on offer. We live in a country where we are used to government supporting Higher Education; now, however, the only source of income to universities for teaching is via student fees, but raising fees is unpopular.  The funding of universities will collapse unless they can either find alternative sources of income, or continue to raise fees in line with inflation, and TEF provides a cover story for doing that.

So we have a system designed to separate winners and losers, but the outcome will depend crucially on two factors: the rate of inflation and the rate of increase in students. The figures in the document have been modelled assuming that the number of students at English Higher Education Institutions will increase at a rate of around 2 per cent per annum (Table 12), and that annual inflation will be around 3 per cent. If either growth in numbers or inflation is lower, then the difference between those who do and don’t get good TEF ratings (and hence the apparent financial benefits of TEF) will decline.

What about the anticipated costs of the TEF?  We are told: “Institutions collectively will experience average annual costs of £22m as a result of familiarising, signing up and applying to the Teaching Excellence Framework, once the TEF covers discipline level assessments. This is equivalent to an average of £53,000 per institution, significantly less than the Research Excellence Framework (REF) at £230,000 per institution per year.” (p. 8). One can only assume that those writing this report have little experience of how academic institutions operate. For instance, they say that “Year One will not represent any additional administrative cost to institutions, as we will use the existing QA process.” I did a quick internet search and immediately found two universities who were advertising now for administrators to work on preparing for the TEF (on salaries of around £30-40K), as well as a consultancy agency that was touting for custom by noting the importance of being “TEF-ready”.

I have yet to get on to the section on costs and benefits of opening the market to ‘alternative providers’…..

If you are concerned at the threats to Higher Education posed by the Bill, please write to your MP - there is a website here that makes it very easy to do so.

Further background reading 
Shaky foundations of the TEF
A lamentable performance by Jo Johnson
More misrepresentation in the Green Paper
The Green Paper’s level playing field risks becoming a morass
NSS and teaching excellence: wrong measure, wrongly analysed
The Higher Education and Research Bill: What's changing?
CDBU's response to the Green Paper
The Alternative White Paper

Tuesday, 24 May 2016

Who wants the TEF?



I'll say this for the White Paper on Higher Education "Success as a Knowledge Economy": it's not as bad as the Green Paper that preceded it. The Green Paper had me abandoning my Christmas shopping for furious tirades against the errors and illogicality that were scattered among the exhausted clichés and management speak (see here, here, here, here and here). So appalled was I at the shoddy standards evident in the Green Paper that I actually went through all the sources quoted in the first section of the White Paper to contact the authors to ask if they were happy with how their work had been reported. I'm pleased to say that out of 12 responses I got, ten were entirely satisfied, and one had just a minor quibble. But what about the twelfth, you ask. What indeed?
When justifying the need for a Teaching Excellence Framework (TEF) last November, Jo Johnson used some extremely dodgy statistical analysis of the National Student Survey to support his case that teaching in some quarters was 'lamentable'. I was pleased to see that this reference was expunged from the White Paper. But that left a motheaten hole in the fabric of the argument: if students aren't dissatisfied, then do we really need a TEF?  One could imagine the civil servants rushing around desperate to find a suitably negative statistic. And so they did, citing the 2015 HEPI-HEA Student Academic Experience Survey as showing that "Many students are dissatisfied with the provision they receive, with over 60% of students feeling that all or some elements of their course are worse than expected and a third of these attributing this to concerns with teaching quality." (p 8, para 5).  The same report is subsequently cited as showing that: ".. applicants are currently poorly-informed about the content and teaching structure of courses, as well as the job prospects they can expect. This can lead to regret: the recent Higher Education Academy (HEA)–Higher Education Policy Institute (HEPI) Student Academic Experience Survey found that over one third of undergraduates in England believe their course represents very poor or poor value for money." The trouble is, both of these quotes again use spin and dodgy statistics.
Let's take the 60% dissatisfaction statistic first. The executive summary of the report stated; "Most students are satisfied with their course, with 87% saying that they are very or fairly satisfied, and only 12% feeling that their course is worse than they expected. However, for those students who feel that their course is worse than expected, or worse in some ways and better than others, the number one reason is not the number of contact hours, the size of classes or any problems with feedback but the lack of effort they themselves put in." So how do we get to 60% dissatisfied? This number is arrived at from the finding that 12% said that their experience had been worse than expected, 49% said that it had been better in some ways and worse in others. So it is literally true that there is dissatisfaction with 'some or all elements', but the presentation of the data is clearly biased to accentuate the negative. One is reminded of Hugh in 'The Thick of It' saying "I did not knowingly not tell the truth".
But it gets worse: As pointed out on the Wonkhe blog, among 'key facts' in a briefing note accompanying the White Paper, the claim was reworded to say over 60% of students said they feel their course is worse than expected. The author of the blogpost referred to this as substantial misrepresentation of the survey. This is serious because it appears that in order to make a political point, the government is spreading falsehoods that could cause reputational damage to Universities.
Moving on to perceptions of 'value for money', there are two reasons for giving this low ratings  - you are paying a reasonable amount for something of poor quality, or you are paying an unreasonable amount for something of good quality. Alex Buckley, one of the authors of the report replied to my query to say that while the numeric data were presented accurately, crucial context was omitted. This made it crystal clear it was the money side of the equation that concerned students. He wrote:
"Figure 11 on page 17 of the 2015 HEPI-HEA survey report shows that students from England (paying £9k) and students from Scotland studying in Scotland (paying no fees) have very different perceptions of value for money. And Figure 12 shows that the perceptions of value for money of students from England plummeted at the time of the increase in fees. Half of 2nd year students from England in 2013 thought they were getting good or very good value for money. In 2014, when 2nd years were paying £9k, that figure was a third. (Other global perceptions of quality - satisfaction etc. - did not change). There is something troubling about the Government citing students' perceptions of value for money as a problem for the sector, when they appear to be substantially determined by Government policy, i.e. the level of fees. The survey suggests that an easy way to improve students' perceptions of the value for money of their degree would be to reduce the level of fees - presumably not the message that the Government is trying to get across."
So do students want the TEF? All the indicators say no. Chris Havergal wrote yesterday in the Times Higher about a report by David Greatbatch and Jane Holland in which students in focus groups gave decidedly lukewarm responses to questions about the usefulness of TEF. Insofar as anyone wants information about teaching quality, they want it at the level of courses rather than institutions, but, as an ONS interim review pointed out, the data is mostly too sparse to reliably differentiate among institutions at the subject level. Meanwhile, the NUS has recommended boycotting the National Student Survey, which forms a key part of the metrics to be used by TEF.
This is all rather rum, given that the government claims its reforms will put students at the heart of higher education. It seems that they have underestimated the intelligence of students, who can see through the weasel words and recognise that the main outcome of all the reforms will be further increases in fees.
It's widely anticipated that fees will rise because of the market competition that the White Paper lauds as a positive stimulus to the sector, and it was clear in the Green Paper that one goal of the reforms was to tie the TEF to a regulatory mechanism that would allow higher fees to be set by those with good TEF scores. Perhaps less widely appreciated is that the plan is for the new Office for Students to be funded largely by subscriptions paid by Higher Education Providers. They will have to find the money somewhere, and the obvious way to raise the cash will be by raising fees. So students will be in the heart of the reforms in the sense that having already endured dramatic rises in fees and loss of the maintenance grant, they will now also be picking up the bill for a new regulatory apparatus whose main function is to satisfy a need for information that they do not want.