Showing posts with label plagiarism. Show all posts
Showing posts with label plagiarism. Show all posts

Saturday, 22 February 2025

IEEE Has a Pseudoscience Problem

Guest post by Solal Pirelli


The IEEE, full name Institute of Electrical and Electronics Engineers, is one of the main scientific publishers in domains related to its name. Many IEEE venues, such as ICSE in software engineering and IROS in robotics, are “top” venues that publish important research. While these are conferences and not journals, computer science and related fields are unusual in that conferences are typically the more prestigious option.

But as I’ve covered before in the case of another big computer science publisher, world-class research can coexist with world-class nonsense. Many not-so-top IEEE venues publish “AI gobbledegook sandwiches”, pointless papers that apply standard machine learning or artificial intelligence to basic data sets resulting in vague predictions supposedly improving on ill-defined baselines.

Unfortunately, bad science published by IEEE isn’t limited to boring applications of boring algorithms to boring data. In this blog post, I’ll present IEEE-published pseudoscience of various kinds, show how this correlates with other problems, and discuss why publishers don’t do enough about it. 

All kinds of quackery 

The IEEE has published numerous new “methods” to help providers or users of pseudoscientific disciplines. Ayurveda is enhanced with a “preprocessing framework” to detect diabetes, a neural network to classify herbs, and even an AI assistant. Astrology is automated with a machine learning model. Myers-Briggs personality type testing is granted another neural network.

Some IEEE papers are at the very fringe of pseudoscience, unconventional even by quack standards. A symposium on antennas and propagation published three papers by the same author on “scientific traditional Chinese medicine”, a variant based on electromagnetism and 5G with “supernatural potential” (see here, here and here). An Indian conference on electronics published four papers (here, here, here and here) by the same first author on “electro-homeopathy”, the brainchild of a 19th century Italian count that an Indian high court called “nothing but quackery” a decade before these papers were published.

Of course, no list of pseudoscience would be complete without perpetual motion. That’s right, the IEEE has published two papers on perpetual motion in 2017 and 2022! How these were not desk-rejected is anyone’s guess.

Even work that is not pseudoscientific in itself can propagate harmful or downright absurd stereotypes. Consider what IEEE-published and supposedly peer-reviewed papers have to say about autism: 

  • “Children with autism require constant care because you never know what will trigger them” (source
  • "If symptoms of autism are detected early, children with autism usually return to normal development after effective medical intervention" (source
  • “A baby born with autism spectrum disorder may have a lower-than-average heart rate. Complete blockage of the heart at birth is rare. Abnormal heart rate leads to heart block. So, there is a high chance of the child's death due to permanent heart blockage at any time.” (source)

Why it matters

One may think that such papers won’t cause harm because they’re unlikely to be read, since they are mostly in unknown venues and unrelated to the IEEE’s domain. While I personally disagree since I believe publishing pseudoscience risks breaking the public’s trust in legitimate research, let me provide a more objective argument. Pseudoscience in papers is heavily correlated with other problematic practices that are more difficult to detect automatically. This makes searching for pseudoscience an effective way to find problematic venues, complementary to existing techniques.

The preprocessing framework to detect diabetes with Ayurveda? In a conference that accepts papers on the same day they are submitted, somehow speeding up the weeks or months usually necessary for proper peer review.

The neural network that classifies Ayurvedic herbs? In a conference that plagiarized its peer review policy from Elsevier’s “Transport Policy” journal. Look for fragments of this policy in your favorite search engine and you’ll find a surprising number of venues that have done so, seemingly without noticing the references to Transport Policy.

The four papers on electro-homeopathy? In a conference that published a mathematical “algorithm” amounting to high school mathematics. While exact definitions of “novelty” vary, no one could credibly claim that this paper is novel enough for a scientific conference.

The 2017 paper on perpetual motion? In a conference that didn’t notice an entirely plagiarized section in that paper, ironically from a source explaining why perpetual motion is impossible. How this is compatible with IEEE’s policy of checking all content for plagiarism is unclear.

The paper claiming “you never know what will trigger” autistic children? In a conference supposedly happening in a London office building, whose four IEEE-published editions only feature one paper from a European university among a sea of India-based authors. Did the authors of this conference’s papers really travel to the other side of the globe to present in a place not designed for presentations?

The neural network for Myers-Briggs? In a conference chaired by a professor whose Russian university is under sanctions from the US, the EU, Ukraine, and even Switzerland!

Action is rare 

The expected process here would be to report this nonsense to the publisher, who would investigate, quickly conclude these papers should never have been published, lose faith in the peer review process that led to their acceptance, and issue retractions. Barring extremely strong evidence from conference chairs that some cases were truly one-off exceptions, such retractions would cover entire editions of conferences.

This happens… sometimes. The IEEE has retracted papers before, such as this one after “only” five months. They have also retracted entire venues, such as this one totaling 400 papers, four years after it was reported.
 
But the IEEE frequently does not react at all to reports. Guillaume Cabanac, who specializes in scientific fraud detection, has repeatedly and publicly called them out. For instance, he’s reported telltale signs of ChatGPT as in this paper that includes “Regenerate Response” in the middle of text and this paper that includes “I am unable to […] due to the fact I am an AI language model”. He’s also reported “tortured phrases”, attempts at avoiding plagiarism detection that instead create nonsense such as “parcel misfortune” instead of “packet loss” in computer networking, in sometimes large concentrations. Cabanac and other sleuths have published “proceedings-level reports” on PubPeer, such as this one, when entire IEEE conferences have problems. None of the examples in this paragraph have led to any public reaction from the IEEE.

The IEEE occasionally issues “expressions of concern”, such one as for this paper over a year after concrete evidence of plagiarism was publicly reported. But expressions of concerns are not retractions. In mid-2023, Retraction Watch noted that hundreds of IEEE papers reported by Guillaume Cabanac and Harvard lecturer Kendra Albert were still up for sale. A year and a half later, that remains the case.

One case noted above is particularly noteworthy in terms of both reputation and IEEE awareness: The “scientific TCM” papers were published in the 2022 and 2023 editions of the “International Symposium on Antennas and Propagation”, a 6-decade-old conference whose 2024 edition boasted the IEEE President as a keynote speaker. Clearly, the IEEE is aware of the venue and its papers. What’s the point in “reporting” them?


Processes are inadequate 

The scale of publishers’ actions is nowhere near the scale of the problem. Creating a new conference or journal does not require that much time if the peer reviewing process is fake. As long as the average time it takes a publisher to retract a venue is higher than the time it takes to create a new venue, there won’t be meaningful progress.

Current publisher processes are designed to correct honest mistakes, not to fight malice. The time it takes to contact authors, wait for their response, wait for them to find original data, and so on is worth it when a single paper has a problem that can be explained by human error. But any such process is a waste of time when a paper contains blatant pseudoscience, has obviously been plagiarized, or uses terminology so bizarre no reviewer could have understood it.

To give an example of scale, here’s a collision of pseudoscience and tortured phrases. The paper on an AI assistant for ayurveda mentioned earlier is in the “2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT)”. Guillaume Cabanac’s Problematic Paper Screener currently lists 185 cases of tortured phrases manually confirmed by Cabanac himself, with another 160 pending assessment. These include “herbal language” instead of natural language, “system getting to know” instead of machine learning, “give-up-to-give-up” instead of end-to-end, and “0.33-celebration” instead of third-party.  

Individually contacting and waiting for hundreds of authors just in case they can explain why their paper talks about 0.33-celebrations isn’t going to cut it. Neither is individually contacting and waiting for dozens of conference editors just in case they can explain why their peer review process didn’t spot this nonsense. 

What can we do?

Given the incentives and processes at play, it’s not surprising to see the IEEE or any other big publisher publish pseudoscience. The authors of the papers mentioned in this post probably didn’t do anything illegal, except maybe for occasional plagiarism of copyrighted content, but nobody has the time and money to sue for such boring violations. This gives publishers a double excuse: they’re not publishing anything illegal, and retractions without a solid legal basis could backfire.

The scientific community needs to ban the “incompetence” defense from authors and stop associating with publishers that can’t be bothered to act quickly enough.  

Authors who publish obvious nonsense should not get a chance to explain themselves or “correct” their paper. 

Publishers make enough money from processing and selling articles. They can defend themselves from occasional lawsuits by angry authors, and they can hire scientific integrity specialists.

When I say “scientific integrity specialist”, that can unfortunately be as simple as “person looking for specific keywords in Google Scholar”. It’s what I did to find pseudoscience, and you can do that too. Report these on PubPeer, directly to publishers, or both. You can also go to the Problematic Paper Screener’s page listing articles that have not been manually assessed yet, and follow the instructions.

Finally, remember that most scientists have no idea this is going on. You can help by publicly calling out problematic papers and lack of action. Ask candidates for governance boards in more democratic publishers like the IEEE what they plan to do about fraud. Discourage institutions, especially public ones in democratic countries, from making blanket deals with publishers.

Tuesday, 9 August 2022

Can systematic reviews help clean up science?

 

The systematic review was not turning out as Lorna had expected

Why do people take the risk of publishing fraudulent papers, when it is easy to detect the fraud? One answer is that they don’t expect to be caught. A consequence of the growth in systematic reviews is that this assumption may no longer be safe. 

In June I participated in a symposium organised by the LMU Open Science Center in Munich entitled “How paper mills publish fake science industrial-style – is there really a problem and how does it work?” The presentations are available here. I focused on the weird phenomenon of papers containing “tortured phrases”, briefly reviewed here. For a fuller account see here. These are fakes that are easy to detect, because, in the course of trying to circumvent plagiarism detection software, they change words, with often unintentionally hilarious consequences. For instance, “breast cancer” becomes “bosom peril” and “random value” becomes “irregular esteem”. Most of these papers make no sense at all – they may include recycled figures from other papers. They are typically highly technical and so to someone without expertise in the area they may seem valid, but anyone familiar with the area will realise that someone who writes “flag to commotion” instead of “signal to noise” is a hoaxer. 

Speakers at the symposium drew attention to other kinds of paper mill whose output is less conspicuously weird. Jennifer Byrne documented industrial-scale research fraud in papers on single gene analyses that were created by templates, and which purported to provide data on under-studied genes in human cancer models. Even an expert in the field may be hoodwinked by these. I addressed the question of “does it matter?” For the nonsense papers generated using tortured phrases, it could be argued that it doesn’t, because nobody will try to build on that research. But there are still victims: authors of these fraudulent papers may outcompete other, honest scientists for jobs and promotion, journals and publishers will suffer reputational damage, and public trust in science is harmed. But what intrigued me was that the authors of these papers may also be regarded as victims, because they will have on public record a paper that is evidently fraudulent. It seems that either they are unaware of just how crazy the paper appears, or that they assume nobody will read it anyway. 

The latter assumption may have been true a couple of decades ago, but with the growth of systematic reviews, researchers are scrutinizing many papers that previously would have been ignored. I was chatting with John Loadsman, who in his role as editor of Anaesthesia and Intensive Care has uncovered numerous cases of fraud. He observed that many paper mill outputs never get read because, just on the basis of the title or abstract, they appear trivial or uninteresting. However, when you do a systematic review, you are supposed to read everything relevant to the research question, and evaluate it, so these odd papers may come to light. 

I’ve previously blogged about the importance of systematic reviews for avoiding cherrypicking of the literature. Of course, evaluation of papers is often done poorly or not at all, in which case the fraudulent papers just pollute the literature when added to a meta-analysis. But I’m intrigued at the idea that systematic reviews might also serve the purpose of putting the spotlight on dodgy science in general, and fraudsters in particular, by forcing us to read things thoroughly. I therefore asked Twitter for examples – I asked specifically about meta-analysis but the responses covered systematic reviews more broadly, and were wide-ranging both in the types of issue that were uncovered and the subject areas. 

Twitter did not disappoint: I received numerous examples – more than I can include here. Much of what was described did not sound like the work of paper mills, but did include fraudulent data manipulation, plagiarism, duplication of data in different papers, and analytic errors. Here are some examples: 

Paper mills and template papers

Jennifer Byrne noted how she became aware of paper mills when looking for studies of a particular gene she was interested in, which was generally under-researched. Two things raised her suspicions: a sudden spike in studies of the gene, plus series of papers that had the same structure, as if constructed from a template. Subsequently, with Cyril LabbĂ©, who developed an automated Seek & Blastn tool to assess nucleotide sequences, she found numerous errors in the reagents and specification of genetic sequences of these repetitive papers, and it became clear that they were fraudulent. 

An example of a systematic review that discovered a startling level of inadequate and possibly fraudulent research was focused on the effect of tranexamic acid on post-partum haemorrhage: out of 26 reports, eight had sections of identical or very similar text, despite apparently coming from different trials. This is similar to what has been described for papers from paper mills, which are constructed from a template. And, as might be expected for a paper mill output, there were also numerous statistical and methodological errors, and some cases without ethical approval. (Thanks to @jd_wilko for pointing me to this example). 

Plagiarism 

Back in 2006, Iain Chalmers, who is generally ahead of his time, noted that systematic reviews could root out cases of plagiarism, citing the example of Asim Kurjak, whose paper on epidural analgesia in labour was heavily plagiarised

Data duplication 

Meta-analysis can throw up cases where the same study is reported in two or more papers, with no indication that this is the same data. Although this might seem like a minor problem compared with fraud, it can be serious, because if the duplication is missed in a meta-analysis, that study will be given more weight than it should have. Ioana Cristea noted that such ‘zombie papers’ have cropped up in a meta-analysis she is currently analysing. 

Tampering with peer review 

When a paper considered for a meta-analysis seems dubious, it raises the question of whether proper peer review procedures were followed. It helps if the journal adopts open peer review. Robin N. Kok reported a paper where the same person was listed as an author and a peer reviewer. This was eventually retracted.  

Data seem too good to be true 

This piece in Science tells the story of Qian Zhang, who published a series of studies on impact of cartoon violence in children which on the one hand had remarkably large samples of children all at the same age, and on the other hand had similar samples across apparently different studies.  Because of their enormous size, Zhang’s papers distorted any meta-analysis they were included in. 

Aaron Charlton cited another case, where serious anomalies were picked up in a study on marketing in the course of a meta-analysis. The paper was ultimately retracted 3 years after the concerns were raised, after defensive responses from some of the authors, challenging the meta-analysts. 

This case flagged by Neil O’Connell is especially useful, as it documents a range of methods used to evaluate suspect research. The dodgy work was first flagged up in a meta-analysis of cognitive behaviour therapy for chronic pain.  Three papers with the same lead author, M. Monticone, obtained results that were discrepant with the rest of the literature, with much bigger effect sizes. The meta-analysts then looked at other trials by the same team and found that there was a 6-fold difference between the lower confidence interval of the Monticone studies and the upper confidence interval of all others combined. The paper also reports email exchanges with Dr Monticone that may be of interest to readers. 

Poor methodology 

Fiona Ramage told me that in the course of doing a preclinical systematic review and meta-analysis of nutritional neuroscience, she encountered numerous errors of basic methodology and statistics, e.g. dozens of papers where error bars were presented without indicating if they show SE or SD; studies claiming differences between groups without a direct statistical comparison. This is more likely to be due to ignorance or honest error than to malpractice, but it needs to be flagged up so that the literature is not polluted by erroneous data.

What are the consequences?

Of course, the potential of systematic reviews to detect bad science is only realised if the dodgy papers are indeed weeded out of the literature, and people who commit scientific fraud are fired. Journals and publishers have started to respond to paper mills, but, as Ivan Oransky has commented, this is a game of Whac-a-Mole, and "the process of retracting a paper remains comically clumsy, slow and opaque”. 

I was surprised that even when confronted with an obvious case of a paper that had both numerous tortured phrases and plagiarism, the response from the publisher was slow – e.g. this comically worded example is still not retracted, even though the publisher’s research integrity office acknowledged my email expressing concern over 2 months ago.  But 2 months is nothing. Guillaume Cabanac recently tweeted about a "barn door" case of plagiarism that has just been retracted 20 years after it was first flagged up.  When I discuss the slow responses to concerns with publishers, they invariably say that they are being kept very busy with a huge volume of material from paper mills. To which I answer, you are making immense profits, so perhaps some could be channeled into employing more people to tackle this problem. As I am fond of pointing out, I regard a publisher who leaves seriously problematic studies in the literature as analogous to a restauranteur that serves poisoned food to customers. 

Publishers may be responsible for correcting the scientific record, but it is institutional employers who need to deal with those who commit malpractice. Many institutions don’t seem to take fraud seriously. This point was made back in 2006 by Iain Chalmers, who described the lenient treatment of Asim Kurjak, and argued for public naming and shaming of those who are found guilty of scientific misconduct. Unfortunately, there’s not much evidence that his advice has been heeded. Consider this recent example of a director of a primate reseach lab who admitted fraud, but is still in post. (Here the fraud was highlighted by a whistleblower rather than a systematic review, but this illustrates the difficulty of tackling fraud when there are only minor consequences for fraudsters). 

Could a move towards "slow science" help? In the humanities, literary scholars pride themselves on “close reading” of texts. In science, we are often so focused on speed and concision, that we tend to lose the ability to focus deeply on a text, especially if it is boring. The practice of doing a systematic review should in principle develop better skills in evaluation of individual papers, and in so doing help cleanse the literature from papers that should never have got published in the first place. John Loadsman has suggested we should not only read papers carefully, but should recalibrate ourselves to have a very high “index of suspicion” rather than embracing the default assumption that everyone is honest. 

P.S

Many thanks to everyone who sent in examples. Sorry I could not include everything. Please feel free to add other examples or reactions in the Comments – these tend to get overwhelmed with adverts for penis enlargement or (ironically) essay mills, and so are moderated, but I do check them and relevant comments will eventually appear.

PPS. Florian Naudet sent a couple of relevant links that readers might enjoy: 

Fascinating article by Fanelli et al who looked at how inclusion of retracted papers affected meta-analyses: https://www.tandfonline.com/doi/full/10.1080/08989621.2021.1947810  

And this piece by Lawrence et al shows the dangers of meta-analyses when there is insufficient scrutiny of the papers that are included: https://www.nature.com/articles/s41591-021-01535-y  

Also, Joseph Lee tweeted about this paper about inclusion of papers from predatory publications in meta-analyses: https://jmla.pitt.edu/ojs/jmla/article/view/491 

PPPS. 11th August 2022

A couple of days after posting this, I received a copy of "Systematic Reviews in Health Research" edited by Egger, Higgins and Davey Smith. Needless to say, the first thing I did was to look up "fraud" in the index. Although there are only a couple of pages on this, the examples are striking. 

First, a study by Nowbar et al (2014) on bone marrow stem cells for heart disease found that in a review of 133 reports, over 600 discrepancies were found, and the number of discrepancies increased with the reported effect size. There's a trail of comments on Pubpeer relating to some of the sources, e.g. https://pubpeer.com/publications/B346354468C121A468D30FDA0E295E.

Another example concerns the use of beta-blockers during surgery. A series of studies from one centre (the DECREASE trials) showing good evidence of effectiveness was investigated and found to be inadequate, with missing data and failure to follow research protocols. When these studies were omitted from a meta-analysis, the conclusion was that, far from receiving benefit from beta-blockers, patients in the treatment group were more likely to die (Bouri et al, 2014). 

 PPPPS, 18th August 2022

This comment by Jennifer Byrne was blocked by Blogger - possibly because it contained weblinks.

Anyhow, here is what she said:

I agree, reading both widely and deeply can help to identify problematic papers, and an ideal time for this to happen is when authors are writing either narrative or systematic reviews. Here's another two examples where Prof Carlo Galli and colleagues identified similar papers that may have been based on templates: https://www.mdpi.com/2304-6775/7/4/67, https://link.springer.com/article/10.1007/s11192-022-04434-2