Showing posts with label higher education. Show all posts
Showing posts with label higher education. Show all posts

Sunday, 9 February 2020

Stemming the flood of illegal external examiners


On 21st January, the Times Higher Education published a short piece about Professor Eric Barendt, an academic lawyer at UCL, who had been told that he had to submit his passport to another University in order to be acceptable as an external examiner. He thought this was preposterous, and declined to do so. The reaction on Twitter indicated that passport checks were now widespread in British universities, and many academics were unhappy about it 

My sympathies are with Prof Barendt, and I've decided that I too will not agree to be an external examiner if I am required to provide my passport to prove I am eligible. In fact, a few days after this story broke, I was invited to be an external examiner, and agreed only on condition that I did not have to provide my passport. Alas, it looks like this means I won't be examining the thesis.

This may look like petulance: refusal to comply with what is not an burdensome requirement creates difficulties for a blameless candidate and their supervisor. So let me explain why I think it is important.

External examining is a highly skilled, high-stakes, onerous task for which one is paid not much more than the minimum wage. The going rate varies from institution to institution, but in my recent experience you may get around £180 to £240. You have to read and evaluate a thesis that represents 3 years' worth of work (around 40,000-50,000 words in my discipline), visit the candidate's home institution to conduct an oral examination that lasts around 2-3 hours, ensure that any corrections are done to your satisfaction, and write a report with recommendations. Nobody does this for the money. Rather, like so much in academia, the whole system survives by a quid pro quo: you know that when your own students need examining, you'll want to find external examiners for them. With a strong student, examining can have its own intrinsic rewards, but it can also be highly stressful if there are problems with the thesis. So overall, all of the academics involved in this process know that the external examiner is doing a favour for another institution by agreeing to take on this extra job.

When I first did examining, many years ago, arrangements for selecting examiners were pretty informal. Times change, and everything has got more official and bureaucratic. Many institutions now require external examiners to provide proof of their competence to do the job (a CV and/or list of previous candidates examined), and some have guidelines to avoid too much chumminess between supervisor and external examiner (no co-authorships, for example). I can see that these requirements, have a point in preserving the integrity of the examination system.

But the passport check is really the last straw. It's senseless on two counts. First, it implies that academic institutions classify external examiners as employees, even though they are doing a one-off task for which the pay is trivial. Second, as Prof Barendt noted, it means that they don't trust other academic institutions to do proper checks of right to work. Now, it may be that there are some dodgy places where this is the case, but it seems reasonable to assume that Higher Education Institutions recognised by the Office for Students will be compliant with the law on this point. What is weird is that when I protest about the passport check for external examiners, some colleagues say, "But if the institution didn't do these checks, they'd be liable for enormous fines". Well, given that is the case, then surely it's safe to assume that the institution that actually employs the external examiner will have done the checks. I can understand that institutions might want an option of conducting checks in rare cases where there was reason to doubt this was true. It's the mandatory nature of the checks that are otiose in 99% of cases that is so exasperating.

Some years ago, in a different context, I wrote a piece about expansion of research regulation in academic life. Many of the points I made there apply to this situation. Bureaucracy creeps up on us by a series of stealthy small steps, until we suddenly find ourselves engulfed by it. Yes, showing a passport is a trivial matter, but I think that if we don't resist this kind of thing, it will only get worse.

P.S. Eric Barendt has pointed me to a piece he wrote on this topic for the Oxford Magazine (2020, No. 416, pp 8-10). I don't think this publication is available online, so here is just a short quote from it concerning the legal aspects of passport checks - something that has been discussed on Twitter in response to this blogpost.
An employer breaks the law only if it employs an illegal immigrant, not because it fails to conduct passport checks. If it is confident it is employing a UK national (or other person with a right to work in the UK such as an EEA or Swiss national), then it has nothing to worry about. So an automatic request is unnecessary. It reveals what may be termed a culture of ‘over-compliance’ with government policy. Of course, it is a sensible, indeed a vital, step to take, if a university, or indeed any employer, has doubts about the immigration status of anyone it is contemplating employing, but common sense surely suggests it is quite unwarranted when it engages someone whom it ought to trust.
Another issue tackled by Eric's piece is whether it is reasonable for Universities to treat external examiners as employees:
... it is hard to see why an external examiner, particularly of a doctoral thesis, should be treated as an employee of the host university, when an academic reviewer of a book proposal is not regarded as an employee of the publisher which engaged him (or her) to review it.
It has been suggested on Twitter that if we are to be regarded as employees, we should be paid an appropriate wage, and the post should be advertised!  

Sunday, 3 March 2019

Benchmarking in the TEF: Something doesn't add up (v.2)





Update: March 6th:   
This is version 2 of this blogpost, taking into account new insights into the weird z-scores used in TEF.  I had originally suggested there might be an algebraic error in the formula used to derive z-scores: I now realise there is a simpler explanation, which is that the z-scores used in TEF are not calculated in the usual way, with the standard deviation as denominator, but rather with the standard error of measurement as denominator. 
In exploring this issue, I've greatly benefited from working openly with a R markdown script on Github, as that has allowed others with statistical expertise to propose alternative analyses and explanations. This process is continuing, and those interested in technical details can follow developments as they happen on Github, see benchmarking_Feb2019.rmd.
Maybe my experience will encourage OfS to adopt reproducible working practices.


I'm a long-term critic of the Teaching Excellence and Student Outcomes Framework (TEF). I've put forward a swathe of arguments against the rationale for TEF in this lecture, as well as blogging for the Council for Defence of British Universities (CDBU) about problems with its rationale and statistical methods. But this week, things got even more interesting. In poking about in the data behind the TEF, I stumbled upon some anomalies that suggest to me that the TEF is not just misguided, but also is based on a foundation of statistical error.

Statistical critiques of TEF are not new. This week, the Royal Statistical Society wrote a scathing report on the statistical limitations of TEF, complaining that their previous evidence to TEF evaluations had been ignored, and stating: 'We are extremely worried about the entire benchmarking concept and implementation. It is at the heart of TEF and has an inordinately large influence on the final TEF outcome'. They expressed particular concern about the lack of clarity regarding the benchmarking methodology, which made it impossible to check results.

This reflects concerns I have had, which have led me to do further analyses of the publicly available TEF datasets. The conclusion I have come to is that the way in which z-scores are defined is very different from the usual interpretation, and leads to massive overdiagnosis of under- and over-performing institutions.

Needless, to say, this is all quite technical, but even if you don't follow the maths, I suggest you just consider the analyses reported below, in which I compare the benchmarking output from the Draper and Gittoes method with that from an alternative approach.

Draper & Gittoes (2004): a toy example

Benchmarking is intended to provide a way of comparing institutions on some metric, while taking into account differences between institutions in characteristics that might be expected to affect their performance, such as the subjects of study, and the social backgrounds of students. I will refer to these as 'contextual factors'.

The method used to do benchmarking comes from Draper and Gittoes, 2004, and is explained in this document by the Higher Education Statistics Agency: HESA. A further discussion of the method can be found in this pdf of slides from a talk by Draper (2006).

Draper (2006) provides a 'small world' example with 5 universities and 2 binary contextual categories, age and gender, to yield four combinations of contextual factors. The numbers in the top part of the chart are the proportions in each contextual (PCF) category meeting the criterion of student continuation.  The numbers in the bottom part are the numbers of students in each contextual category.


Table 1. Small world example from Draper 2006, showing % passing benchmark (top) and N students (bottom)

Essentially, the obtained score (weighted mean column) for an institution is an average of indicator values for each combination of contextual factors, weighted by the numbers with each combination of contextual factors in the institution. The benchmarked score is computed by taking the average score for each combination across all institutions (bottom row of top table) and then for each institution creating a mean score, weighted by the number in each category for a that institution. Though cumbersome (and hard to explain in words!) it is not difficult to compute.  You can find an R markdown script that does the computation here (see benchmarking_Feb2019.rmd, benchmark_function). The difference between obtained values and benchmarked value can then be computed, to see if the institution is scoring above expectation (positive difference) or below expectation (negative difference).  Results for the small world example are shown in Table 2.
Table 2. Benchmarks (Ei) computed for small world example
The column headed Oi is the observed proportion with a pass mark on the indicator (student continuation), Ei is the benchmark (expected) value for each institution, and Di is the difference between the two.

Computing standard errors of difference scores

The next step is far more complex. A z-score is computed by dividing the difference between observed and expected values on an indicator (Di) by a denominator, which is variously referred to as a standard deviation and a standard error in the documents on benchmarking.

For those who are not trained in statistics, the basic logic here is that the estimate of an institution's performance will be more labile if it is based on a small sample. If the institution takes on only 5 students each year, then estimates of completion rates from year to year will be variable - in a year where one student drops out, then the completion rate is only 80%, but if none drop out it will be 100%. You would not expect it to be constant because of random factors outside the control of the institution will affect student drop-outs. In contrast, for an institution with 1000 students, we will see much less variation from year to year. The standard error provides an estimate of the extent to which we expect the estimate of average drop-out to vary from year to year, taking size of population into account. 

To interpret benchmarked scores we need a way of estimating the standard error of the difference between the observed score on a metric (such as completion rate) and the benchmarked score, reflecting how much we would expect this to vary from one occasion to another. Only then can we judge whether the institution's performance is in line with expectation. 

Draper (2006) walks the reader through a standard method for computing the standard errors, based on the rather daunting formulae of Figure 1. The values in the SE column of table 2 are computed this way, and the z-scores are obtained by dividing each Di value by its corresponding SE.

Fomulae 5 to 8 are used to compute difference scores and standard errors (Draper, 2006)
Now anyone familiar with z-scores will notice something pretty odd about the values in Table 2. The absolute z-scores given by this method seem remarkably large: In this toy example, we see z-scores with absolute values of 5, 9 and 20.  Usually z-scores range from about -3 to 3. (Draper noted this point).


Z-scores in real TEF data

Next, I downloaded some real TEF data, so I could see whether the distribution of z-scores was unusual. Data from Year 2 (2017) in .csv format can be downloaded from this website.
The z-scores here have been computed by HESA. Here is the distribution of core z-scores for one of the metrics (Non-continuation) for the 233 institutions with data on FullTime students.

The distribution is completely out of line with what we would expect from a z-score distribution.  Absolute z-scores greater than 3, which should be vanishingly rare, are common - with the exact number varying across the six available metrics, but ranging from 33% to 58%.

Yet, they are interpreted in TEF as if a large z-score is an indicator of abnormally good or poor performance:

From p. 42  of this pdf giving Technical Specifications:

"In TEF metrics the number of standard deviations that the indicator is from the benchmark is given as the Z-score. Differences from a benchmark with a Z-score +/-1.9623 will be considered statistically significant. This is equivalent to a 95% confidence interval (that is, we can have 95% confidence that the difference is not due to chance)."

What does the z-score represent?

Z-scores feature heavily in my line of work: in psychological assessment they are used to identify people whose problems are outside the normal range. However, they aren't computed like the TEF z-scores, because they involve dividing a mean score by the standard deviation, rather than by the standard error.


It's easiest to explain this by an analogy. I'm 169 cm tall. Suppose you want to find out if that's out of line with the population of women in Oxford. You measure 10,000 women and find their mean height is 170 cm, with a standard deviation of 3. On a conventional z-score, my height is unremarkable. You just divide the difference between my height and the population height and divide by the standard deviation, -1/3, to give a z-score of -0.33. That's well within the normal limits used by TEF of -1.96 to 1.96.

Now let's compute the standard error of the population mean - to do that we compute the standard error, which is the standard deviation divided by the square root of the sample size, which gives 3/100 or .03. From that information we can get an estimate of the precision of our estimate of the population mean: we multiply the SE by 1.96, and add and subtract that value to the mean to get 95% confidence limits, which are 169.94 and 170.06. If we were to compute the z-score corresponding to my height using the SE instead of the SD, I would seem to be alarmingly short: the value would be -1/.03 = -33.33.

So what does that mean? Well the second z-score based on the SE does not test whether my height is in line with the population of 10,000 women. It tests whether my height can be regarded as equivalent to that of the average from that population. Because the population is very large, the estimate of the average is very precise, and my height is outside the error of measurement for the mean.

The problem with the TEF data is that they use the latter, SE-based method to evaluate differences from the benchmark value, but appear to interpret it as if it was a conventional SD-based z-score:

E.g. in the Technical Specificiations document (5.63):

As a test of the likelihood that a difference between a provider’s benchmark and its indicator is due to chance alone, a z-score +/- 3.0 means the likelihood of the difference being due to chance alone has reduced substantially and is negligible.

As illustrated with the height analogy, the SE-based method seems designed to over-identify high and low-achieving institutions. The only step taken to counteract this trend is an ad hoc one: because large institutions are particularly prone to obtain extreme scores, a large absolute z-score is only flagged as 'significant' if the absolute difference score is also greater than 2 or 3 percentage points. Nevertheless, the number of flagged institutions for each metric, is still far higher than would be the case if conventional z-scores based on the SD were used.

Relationship between SE-based and SD-based z-scores
(N.B. Thanks to Jonathan Mellon who noted an error in my script for computing the true z-scores. 
This update and correction made 20.20 p.m. on 6 March 2019).

I computed conventional z-scores by dividing each institution's difference from benchmark by the SD of for difference scores for all institutions and plotted it against the TEF z-scores. An example for one of the metrics is shown below. The range is in line with expectation (most values between -3 and +3) for the conventional z-scores, but much bigger for the TEF z-scores.




Conversion of z-scores into flags

In TEF benchmarking, TEF z-scores are converted into 'flags', ranging from - - or -, to denote performance below expectation, up to + or ++ for performance above expectation, with = used to indicate performance in line with expectation. It is these flags that the TEF panel considers when deciding which award (Gold, Silver or Bronze) to award.

Draper-Gittoes z-scores are flagged for significance as follows:
  •  - - z-score of -3 or less, AND an absolute difference between observed and expected values of 3%. 
  •  - z-score of -2 or less, AND an absolute difference between observed and expected values of 2%. 
  •  + z-score of 2 or more, AND an absolute difference between observed and expected values of 2%.   
  • ++ z-score of 3 or more, AND an absolute difference between observed and expected values of 3%. 
Given the problems with the method outlined above, this method is likely to massively overdiagnose both problems and good performance.

Using quantiles rather than TEF z-scores

Given that the z-scores obtained with the Draper-Gittoes method are so extreme, it could be argued that flags should be based on quantiles rather than z-score cutoffs, omitting the additional absolute difference criterion. For instance, for the Year 2 TEF data (Core z-scores) we can find cutoffs corresponding to the most extreme 5% or 1%.  If flags were based on these, then we would award extreme flags (- - or ++) only to those with negative z-scores of -13.7 or less, or positive score of 14.6 or more; less extreme flags would be awarded to those with negative z-score of -7 or less (- flag), or positive z-score of 8.6 or more (+).

Update 6th March: An alternative way of achieving the same end would be to use the TEF cutoffs with conventional z-scores; this would achieve a very similar result.

Deriving award levels from flags

It is interesting to consider how this change in procedure would affect the allocation of awards. In TEF, the mapping from raw data to awards is complex and involves more than just a consideration of flags: qualitative information is also taken into account. Furthermore, as well as the core metrics, which we have looked at above, the split metrics are also considered - i.e. flags are also awarded for subcategories, such as male/female, disabled/non-disabled: in all there are around 130 flags awarded across the six metrics for each institution. But not all flags are treated equally: the three metrics based on the National Student Survey are given half the weight of other metrics.

Not surprisingly, if we were to recompute flag scores based on quantiles, rather than using the computed z-scores, the proportions of institutions with Bronze or Gold awards drops massively.

When TEF awards were first announced, there was a great deal of publicity around the award of Bronze to certain high-profile institutions, in particularly the London School of Economics, Southampton University, University of Liverpool and the School of Oriental and African Studies. On the basis of quantile scores for Core metrics, none of these would meet criteria for Bronze: their flag scores would be -1, 0, -.5 and 0 respectively. But these are not the only institutions to see a change in award when quantiles are used. The majority of smaller institutions awarded Bronze obtain flag scores of zero.

The same is true of Gold Awards. Most institutions that were deemed to significantly outperform their benchmarks no longer do so if quantiles are used.

Conclusion

Should we therefore change the criteria used in benchmarking and adopt quantile scores? Because I think there are other conceptual problems with benchmarking, and indeed with TEF in general, I would not make that recommendation. I would prefer to see TEF abandoned. I hope the current analysis can at least draw people's attention to the questionable use of statistics used in deriving z-scores and their corresponding flags. The difference between a Bronze, Silver and Gold can potentially have a large impact on an institution's reputation. The current system for allocating these awards is not, to my mind, defensible.

I will, of course, be receptive to attempts to defend it or to show errors in my analysis, which is fully documented with scripts on github, benchmarking_Feb2019.rmd.


Saturday, 11 August 2018

More haste less speed in calls for grant proposals


Helpful advice from the World Bank

This blogpost was prompted by a funding call announced this week by the Economic and Social Research Council (ESRC)  , which included the following key dates:
  • Opening date for proposals – 6 August 2018 
  • Closing date for proposals – 18 September 2018 
  • PI response invited – 23 October 2018 
  • PI response due – 29 October 2018 
  • Panel – 3 December 2018 
  • Grants start – 14 February 2019 
As pointed out by Adam Golberg (@cash4questions), Research Development Manager at Nottingham University, on Twitter, this is very short notice to prepare an application for substantial funding:
I make this about 30 working days notice. For a call issued in August. For projects of 36 months, up to £900k - substantial, for social sciences. With only one bid allowed to be led from each institution, so likely requiring an internal sift. 

I thought it worth raising this with ESRC, and they replied promptly, saying:
To access funds for this call we’ve had to adhere to a very tight spending timeframe. We’ve had to balance the call opening time with a robust peer review process and a Feb 2019 project start. We know this is a challenge, but it was a now or never funding opportunity for us.
 
They suggested I email them for more information, and I’ve done that, so will update this post if I hear more. I’m particularly curious about what is the reason for the tight spending timeframe and the inflexible February 2019 start.

This exchange led to discussion on Twitter which I have gathered together here.

It’s clear that from the responses that this kind of time-frame is not unusual, and I have been sent some other examples. For instance this ESRC Leadership Fellowship (£100,000 for 12 months) had a call for proposals issued on 16th November 2017, with a deadline for submissions of 3 January. When you factor in that most universities shut down from late December until early January, and so this would need to be with administrators before the Christmas break, this gives applicants around 30 days to construct a competitive proposal. But it’s not only ESRC that does this, and I am less interested in pointing the finger at a particular funder – who may well be working under pressures outside their control - than just raising the issue of why this needs a rethink. I see five problems with these short lead times:

1. Poorer quality of proposals 
The most obvious problem is that a hastily written proposal is likely to be weaker than one that is given more detailed consideration. The only good thing you might say about the time pressure is that it is likely to reduce the number of proposals, which reduces the load on the funder’s administration. It’s not clear, however, whether this is an intended consequence.

2. Stress on academic staff 
There is ample evidence that academic staff in the UK have high stress levels, often linked to a sense of increasing demands and high workload. A good academic shows high attention to detail and is at pains to get things right: research is not something that can be done well under tight time pressure. So holding up the offer of a large grant with only a short time period to prepare a proposal is bound to increase stress: do you drop everything else to focus on grant-writing, or pass by the opportunity to enter the competition?

Where the interval between the funding call and the deadline occurs over a holiday period, some might find this beneficial, as other demands such as teaching are lower. But many people plan to take a vacation, and should be able to have a complete escape from work for at least a week or two. Others will have scheduled the time for preparing lectures, doing research, or writing papers. Having to defer those activities in order to meet a tight deadline just induces more sense of overload and guilt at having a growing backlog of work.

3. Equity issues 
These points about vacations are particularly pertinent for those with children at home during the holidays, as pointed out in a series of tweets by Melissa Terras, Professor of Digital Cultural Heritage at Edinburgh University, who said:
I complained once to the AHRC about a call announced in November with a closing date of early January - giving people the chance to work over the Xmas shutdown on it. I wasn't applying to the call myself, but pointed out that it meant people with - say - school age kids - wouldn't have a "clear" Xmas shutdown to work on it, so it was prejudice against that cohort. They listened, apologised, and extended the deadline for a month, which I was thankful for. But we shouldn't have to explain this to them. Have RCUK done their implicit bias training?

4. Stress on administrative staff 
One person who contacted me via email pointed out that many funders, including ESRC, ask institutions to filter out uncompetitive proposals through internal review. That could mean senior research administrators organising exploratory workshops, soliciting input from potential PIs, having people present their ideas, and considering collaborations with other institutions. None of that will be possible in a 30-day time frame. And for the administrators who do the routine work of checking grants for accuracy of funding bids and compliance with university and funder requirements, I suspect it’s not unusual to be dealing with a stressed researcher who expects them to do all of this with rapid turnaround, but where the funding scheme virtually guarantees everything is done in a rush, this just gets worse.

5. Perception of unfairness 
Adding in to this toxic mix, we have the possibility of diminished trust in the funding process. My own interest in this issues stems from a time a few years ago when there was a funding call for a rather specific project in my area. The call came just before Christmas, with a deadline in mid January. I had a postdoc who was interested in applying, but after discussing it, we decided not to put in a bid. Part of the reason was that we had both planned a bit of time off over Christmas, but in addition I was suspicious about the combination of short time-scale and specific topic. This made me wonder whether a decision had already been made about who to award the funds to, and the exercise was just to fulfil requirements and give an illusion of fairness and transparency.

Responses on Twitter again indicate that others have had similar concerns. For instance, Jon May, Professor in Psychology at the University of Plymouth, wrote:
I suspect these short deadline calls follow ‘sandboxes’ where a favoured person has invited their (i.e his) friends to pitch ideas for the call. Favoured person cannot bid but friends can and have written the call.
 
And an anonymous correspondent on email noted:
I think unfairness (or the perception of unfairness) is really dangerous – a lot of people I talk to either suspect a stitch-up in terms of who gets the money, or an uneven playing field in terms of who knew this was coming.

So what’s the solution? One option would be to insist that, at least for those dispensing public money, there should be a minimum time between a call for proposals and the submission date: about 3 months would seem reasonable to me.

Comments will be open on this post for a limited time (2 months, since we are in holiday season!) so please add your thoughts.

P.S. Just as I was about to upload this blogpost, I was alerted on Twitter to this call from the World Bank, which is a beautiful illustration of point 5 - if you weren't already well aware this was coming, there would be no hope of applying. Apparently, this is not a 'grant' but a 'contract', but the same problems noted above would apply. The website is dated 2nd August, the closing date is 15th August. There is reference to a webinar for applicants dated 9th July, so presumably some information has been previously circulated, but still with a remarkably short time lag, given that there need to be at least two collaborating institutions (including middle- and low-income countries)
, with letters of support from all collaborators and all end users. Oh, and you are advised ‘Please do not wait until the last minute to submit your proposal’.


Update: 17th August 2018
An ESRC spokesperson sent this reply to my query:

Thank you for getting in touch with us with your concerns about the short call opening time for the recently announced Management Practices and Employee Engagement call, and the fact that it has opened in August.

We welcome feedback from our community on the administration of funding programmes, and we will think carefully about how to respond to these concerns as we design and plan future programmes.

To provide some background to this call. It builds on an open-invite scoping workshop we held in February 2018, at which we sought input from the academic, policy and third-sector communities on the shape of a (then) potential research investment on management practices and employee engagement. We subsequently flagged the likelihood of a funding call around the topic area this summer, both at the scoping workshop itself, as well as in our ongoing engagements with the academic community.

We do our best to make sure that calls are open for as long as possible. We have to balance call opening times with a robust and appropriately timetabled peer review process, feasible project start dates, the right safeguards and compliances, and, in certain cases such as this one, a requirement to spend funds within the financial year. 

We take the concerns that you raise in your email and in your blog post of 11 August 2018 extremely seriously. The high standard of the UK's research is a result of the work of our academic community, and we are committed to delivering a system that respects and responds to their needs. As part of this, we are actively looking into ways to build in longer call lead times and/or pre-announcements of funding opportunities for potential future managed calls in this and other areas.

I would also like to stress that applicants can still submit proposals on the topic of management practices and employee engagement through our standard research grant process, which is open all year round. The peer review system and the Grant Assessment Panel does not take into account the fact that a managed call is open on a topic when awarding funding: decisions are taken based on the excellence of the proposal.

Update: 23rd August 2018
A spokesperson for the World Bank has written to note that the grant scheme alluded to in my postscript did in fact have a 2 month period between the call and submission date. I have apologised to them for suggesting it was shorter than this, and also apologise to readers for providing misleading information. The duration still seems short to me for a call of this nature, but my case is clearly not helped by providing wrong information, and I should have taken greater care to check details. Text of the response from the World Bank is below:
 
We noticed with some concern that in your Aug. 11 blog post, you had singled out a World Bank call for proposals as a “beautiful illustration” of a type of funding call that appears designed to favor an inside candidate. This characterization is entirely inaccurate and appears based on a misperception of the time lag between the announcement of the proposal and the deadline.
Your reference to the 2018 Call for Proposals for Collaborative Data Innovations for Sustainable Development by the World Bank and the Global Partnership for Sustainable Development Data as undermining faith in the funding process seems based on the mistaken assumption that the call was issued on or about August 2. It was not.
The call was announced June 19 on the websites of the World Bank and the GPSDD. This was two months before the closing date, a period we have deemed fair to applicants but also appropriate given our own time constraints. An online seminar was offered to assist prospective applicants, as you note, on July 9.
The seminar drew 127 attendees for whom we provided answers to 147 questions. We are still reviewing submissions for the most recent call for proposals for this project, but our call for the 2017 version elicited 228 proposals, of which 195 met criteria for external review.
As the response to the seminar and the record of submissions indicate, this funding call has been widely seen and provided numerous applicants the opportunity to respond.  To suggest that this has not been an open and fair process does not do it justice.

Here are the links with the announcement dates of June 19th

Saturday, 6 August 2016

Alternative providers and alternative medicine

Jo Johnson thinks that the market in Higher Education is unfair. There are currently stringent procedures in place for any institution that wants to award degrees and call itself a University. One facet of this is the requirement that any new provider must initially have its degrees validated by an established University; The University of Suffolk, which gained University title this month, provides an example of how this can work, but many of those seeking to enter the market are unhappy with current arrangements.  Roxanne Stockwell, Principal of Pearson College, complained: “Until an institution has its own degree-awarding powers, it cannot offer degrees without being validated by an existing university. Under the current system, a new partner has to find a willing validating partner, and it is locked out if it cannot.”

Jo Johnson, in his role as Minister for Universities and Science last year, criticised this arrangement. He memorably said: “I know some validation relationships work well, but the requirement for new providers to seek out a suitable validating body from amongst the pool of incumbents is quite frankly anti-competitive. It’s akin to Byron Burger having to ask permission of McDonald’s to open up a new restaurant. .....I can announce that we will shortly be lifting the moratorium that has been in place for applications for new Degree Awarding Powers and for University Title. Once again, we are opening the doors to new entrants and challenger institutions, all in the interest of increasing the choices available to students.”

So how wide should the door be opened? This raises deep questions about what constitutes a University. Currently, British Universities have a strong international reputation. Historically, they have been subject to strict scrutiny in return for receiving government funding. They combine research and teaching to push forward the boundaries of knowledge, and have trained students to value knowledge for its own sake, not just as a means to an end.

Not everyone wants that kind of education: some students do not enjoy formal academic study and may prefer vocational courses or apprenticeships. It is important that our higher education system caters for them. The question confronting us now is how far we should extend the definition of a University. It is interesting to consider the Word Cloud I made of ‘alternative providers’, shown in Figure 1.

Figure 1: Alternative providers from http://www.hefce.ac.uk/reg/register/getthedata/. Key: Blue have University Title; Brown have Degree-Awarding Powers; Pink offer designated courses; Violet deliver HE as a franchise only. Taken from http://cdbu.org.uk/speedy-entrances-and-sharp-exits-letting-in-more-alternative-providers/.
Among those listed are six institutions providing training in various forms of complementary and alternative medicine. Two of these had progressed to the point of having degree-awarding powers, which is one step below having University Title.

Let us be clear: the subject matter that these institutions teach is not endorsed by serious scientists.  David Colquhoun highlighted a worrying trend for UK Universities to give degrees in ‘anti-science’ back in 2007. In Australia, where chiropractic has been taught within the regular University system, it has come under increasing attack by scientists, who note that by awarding degrees in these subjects, one gives credibility to procedures that are placebos at best and dangerous at worst.

Are these concerns just a sign of anti-competitiveness and elitism? Are scientists trying to squeeze out the alternative providers because they think they’ll poach students from them, just as McDonalds might try to ensure that Byron Burgers are denied space for development? I’d argue not, and furthermore that Jo Johnson’s analogy shows a startling lack of understanding of what a University is all about. Medicine has had many false leads and it would be the height of arrogance to assume that what we know now is the only truth. But the difference between medicine and alternative medicine is not just that there is evidence for effectiveness for medicine; it’s also that in medicine there is a continuous movement to develop and improve theory and practice, rigorously testing and debating ideas and using scientific methods to evaluate them. It’s difficult to do this well and it often goes wrong, but there is broad agreement about the importance of evidence.

Well, you might say, what about religion, another topic that features heavily among alternative providers? Should theology be banned from Universities because it is not evidence-based? The answer is no, and for similar reasons to those given above: the difference between our traditional Universities and the new providers is that Universities teaching theology consider a range of perspectives and teach students critical thinking. You do not need to be a believer in any God to study theology. In contrast, new providers are often narrow in their focus: many of those with religious affiliations look as if they train students in one religious viewpoint and one only.

Defining the difference between what is suitable and not suitable for inclusion in a University degree course is itself an interesting intellectual exercise. We should not assume that something is good just because it already exists, and is bad if it is new.  But if we just ignore this distinction and have a free-for-all whereby anything can be regarded as higher education provided that there are students willing to pay for it, we will end up with a system in which the terms University and Degree will count for very little, and where the survival of a higher education institution has more to do with its marketing skills than its academic standing. In the past, there were few institutions clamouring to become Universities because it was not easy to make a profit from Higher Education. That has all changed now that higher education providers can get their hands on money from the Student Loans Company. Experience to date suggests we need to have more, rather than less, scrutiny of alternative providers in the current financial climate.

One final point: the alternative providers I have discussed here could do very well on the Teaching Excellence Framework (TEF), which will rate higher education institutions according to three main criteria: student satisfaction, drop-out rates and employability. Indeed, if they recruit their students from among disadvantaged social groups, they might well achieve higher TEF scores than more selective institutions, because benchmarking is used to adjust the outcomes. So under Jo Johnson’s oversight, we could end up with a situation where the quality of teaching at the Anglo-European College of Chiropractic is deemed superior to that at the Universities of Oxford and Cambridge. An interesting thought.

Sunday, 17 July 2016

Cost-benefit analysis of the Teaching Excellence Framework

©CartoonStock.com
The government’s new Higher Education and Research Bill gets its second reading this week. One complaint is that it has been rushed in without adequate scrutiny of some key components. I was interested, therefore, to discover, that a Detailed Impact Assessment was published in June, specifically to look at the costs and benefits of the various components of the Bill. What I found was quite shocking: we were being told that the financial benefits of the new Teaching Excellence Framework (TEF) vastly outweighed its costs – yet look in detail and this is all smoke and mirrors.

In particular, the report shows that while the costs of TEF to the higher education sector (confusingly described as ‘business’) are estimated at £20 million, the direct benefits will come to £1,146 million, giving a net benefit of £1,126 million (Table 1). How could the introduction of a new bureaucratic evaluation exercise be so remarkably beneficial? I read on with bated breath.

Well, sad to relate, it’s voodoo analysis.  This becomes clear if you press on to Table 12, which shows the crucial data from statistical modelling. Quite simply, the TEF generates money for institutions that get a good rating because it allows them to increase fees in line with inflation. Institutions that don’t participate in the TEF, or those that fail to get a good enough rating, will not be able to exceed the current 9K per annum fee, and so in real terms their income will decline over time. As far as I can make out, they are not included in Table 1. Furthermore, the increases for the compliant, successful institutions are measured relative to how they would have done if they had not been allowed to raise fees.

So to sum up:
  • You don’t need the TEF to achieve this result. You could get the same outcome by just allowing all institutions to raise fees in line with inflation.
  • As noted in the briefing to the Bill by the House of Commons: “the Bill is expected to result in a net financial benefit to higher education providers of around £1.1billion a year. This is in very large part due to the higher fees that providers with successful TEF outcomes will be able to charge students.” (p. 59)
  • The system is designed for there to be winners and losers, and the losers will inevitably see their real income falling further and further behind the winners, unless inflation is zero.
The impact assessment does consider other options, including that of allowing fee increases in line with inflation provided the institution has a satisfactory Quality Assurance rating. This is rejected on the grounds that: “whilst QA is a good starting point, reliance on QA alone and in the longer-term will not enable significant differentiation of teaching quality to help inform student decisions and encourage institutions to improve their teaching quality.” (p. 37).  This makes clear that one consequence (and one suspects one purpose) of TEF is to facilitate the division into institutional sheep and goats, followed by starvation of the goats.

Another option, which was strongly recommended by many of those who responded to the consultation exercise on the Green Paper which preceded the bill, is to remove the link between TEF and fees. In other words, have some kind of teaching evaluation, where the motivation for taking part would be reputational rather than financial.  This too is rejected as not sufficiently powerful an incentive: “the Research Excellence Framework allocates £1.5bn a year to institutions. To achieve parity of esteem and focus between teaching and research the TEF will need to have a similar level of financial implications.” However, this is rather disingenuous. There is no pot of money on offer. We live in a country where we are used to government supporting Higher Education; now, however, the only source of income to universities for teaching is via student fees, but raising fees is unpopular.  The funding of universities will collapse unless they can either find alternative sources of income, or continue to raise fees in line with inflation, and TEF provides a cover story for doing that.

So we have a system designed to separate winners and losers, but the outcome will depend crucially on two factors: the rate of inflation and the rate of increase in students. The figures in the document have been modelled assuming that the number of students at English Higher Education Institutions will increase at a rate of around 2 per cent per annum (Table 12), and that annual inflation will be around 3 per cent. If either growth in numbers or inflation is lower, then the difference between those who do and don’t get good TEF ratings (and hence the apparent financial benefits of TEF) will decline.

What about the anticipated costs of the TEF?  We are told: “Institutions collectively will experience average annual costs of £22m as a result of familiarising, signing up and applying to the Teaching Excellence Framework, once the TEF covers discipline level assessments. This is equivalent to an average of £53,000 per institution, significantly less than the Research Excellence Framework (REF) at £230,000 per institution per year.” (p. 8). One can only assume that those writing this report have little experience of how academic institutions operate. For instance, they say that “Year One will not represent any additional administrative cost to institutions, as we will use the existing QA process.” I did a quick internet search and immediately found two universities who were advertising now for administrators to work on preparing for the TEF (on salaries of around £30-40K), as well as a consultancy agency that was touting for custom by noting the importance of being “TEF-ready”.

I have yet to get on to the section on costs and benefits of opening the market to ‘alternative providers’…..

If you are concerned at the threats to Higher Education posed by the Bill, please write to your MP - there is a website here that makes it very easy to do so.

Further background reading 
Shaky foundations of the TEF
A lamentable performance by Jo Johnson
More misrepresentation in the Green Paper
The Green Paper’s level playing field risks becoming a morass
NSS and teaching excellence: wrong measure, wrongly analysed
The Higher Education and Research Bill: What's changing?
CDBU's response to the Green Paper
The Alternative White Paper