Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Sunday, 5 July 2026

How paper mills go straight (without going clean): a shift to dual-use services

Guest post by Matt Spick


                           Photo: Parsadanov/Shutterstock.com

From relative obscurity, paper mills have recently moved into the spotlight of academic attention. These organisations - which sell manuscripts or citations to authors to enhance scholarly metrics - are growing so rapidly that many fields are being overwhelmed. The overall proportion of paper mill outputs was estimated at 1.5–2% of all scientific papers published in 2022, but a recent AI-screening cancer study estimated that around 10% of recent cancer manuscripts in some venues may be paper mill products, and in data-intensive fields paper mill outputs can now outnumber legitimate publications. The increased attention has also been driven by the integrity community highlighting unethical behaviours, notably in the FoSci Report 2026, and has resulted in initiatives such as United2Act. This is in addition to publishers and third parties setting up a growing number of integrity checking systems, whether for citation anomalies, duplicated images, or tortured phrases. But this is an adversarial process, and backward-looking checks will inevitably miss what happens next.

One option for paper mills is to stop selling unethical products, and shift to dual-use services instead. The dual-use business model has a long history. During the US prohibition era, manufacturing moonshine came with severe risks. To avoid these risks, and transfer both criminal intentions (mens rea) and criminal action (actus reus) to the customers, the California Vineyardist Association created a front organisation, Fruit Industries Ltd, to sell Vine-Glo. This innovative product consisted of a block of concentrated grape juice, easily dissolved in water to create a refreshing and entirely legal fruit-based drink. The grape juice block also included an explicit warning to purchasers that “After dissolving the brick in a gallon of water, do not place the liquid in a jug away in the cupboard for twenty days, because then it would turn into wine.” And of course there are other examples of unethical actors shifting into legitimate business, whether in laundry services or waste management businesses. Such operations are hard to police, precisely because they overlap with entirely legal commercial interests.

The paper mill equivalents of Vine-Glo might include proprietary software tools to data dredge and p-hack large open datasets, such as the Global Burden of Disease Study or NHANES. Such tools could be used legitimately, but if videos were created explaining how to mass produce p-hacked papers, this would be claimed as “beyond the developers’ control”. Training courses might emerge that offer primers on scientific writing to complete beginners, starting with an icebreaker at 9:00 am on day one and concluding with submission to a PubMed indexed journal at 5:00 pm on day two. Naturally, acceptance would not be guaranteed, and the trainers would stress that this was purely a training service, not a ‘pay for authorship’ model. Proving that these practices were linked to specific examples of problematic authorship would be impossible at the individual paper level.

The lack of action (sadly, there is no such thing as the science police) is especially frustrating in an industry that is slow to respond even in clear cases of retractions being needed, due to a conservative culture around accusations of wrongdoing, which - partly for good reasons - prefers to tolerate higher levels of unethical outputs rather than inadvertently criticising or punishing scientists. At the same time, the surge in publication volumes (measured in the tens of thousands across biomedical and life sciences research) makes it impossible to conclude that no problematic behaviour is occurring.

Mills turning to dual-use products would also compromise large parts of the integrity community's current toolkit: detection based on image duplication, tortured phrases and citation anomalies would largely fail, because the outputs would represent genuine analyses of real data with real (if trivial) results. In turn, this is likely to lead to prevention having to move upstream, to pre-registration before data release and gatekeeping of access.

Will service mills offering software and training replace the existing paper mills? In practice, both models are likely to continue to exist, if only because hiring, promotion and graduation decisions are all influenced by (or completely dependent upon) publication counts. As long as the traditional model remains profitable (with first-author slots advertised at up to USD 5,600, more than sufficient to grease the wheels through editor bribery elsewhere in the publication chain) and the underlying demand continues, there will be no incentive to leave money on the table. Conventional paper mills may be able to use LLMs and agentic platforms to improve the quality of their outputs (as a different form of adversarial adaptation), making them harder to detect and preserving their value proposition for unethical authors who simply want to purchase authorship. Many potential customers are likely to find the scientific publishing process too arcane, and will be happy to continue to buy an end-to-end service. Service mills seem likely to grow in importance, however, for authors who are more risk averse. At the same time - for expert users - agentic AI platforms for science may trivialise secondary outputs such as reviews and open data analyses, providing a third route to ‘enhancing’ an author’s scholarly record. 

This will present a challenge to the wider scientific community, which has only begun to wake up to the problem of paper mills as manuscript factories. In reality the mills are diversifying their offerings, creating the illusion of ‘going straight’ through dual-use products, and it seems inevitable that some expert users will be able to disintermediate the paper mills completely through agentic AI. We need to recognise that all these types of output can dilute the scholarly record, decreasing the signal to noise ratio and delaying the translation of literature into real impact, to the disadvantage of both direct users and also society as a whole.

Note: Comments are welcome, but are moderated, so there may be a delay before they appear. In general, comments are accepted if on-topic and non-anonymous.

 

 


Friday, 21 November 2025

The dangers of using bibliometrics with polluted data

 

This week I attended a Webinar organised by Clarivate on the topic of “Eugene Garfield Centenary Celebration: Past, Present and Future of Scientometrics”. This covered the early history of the late Eugene Garfield’s work as well as current developments and future trends. The historical sessions were fascinating, describing the remarkable innovations that Garfield made in his quest to understand the body of scientific information as a network. He realised that similarities between articles could be identified by shared citations, and, in the 1950s he devised systems for capturing this information using punched cards. I am old enough to remember going to the library in the 1970s to consult Science Citation Index, which could not only point me to important articles in my field, but often led me in strange directions as I stumbled upon other fascinating topics. 

Garfield is known as the father of the Journal Impact Factor, which is regarded by many as an abomination that distorts publishing behaviour because of its connotations of prestige. It was, however, originally conceived as an index that would help librarians decide which journals to purchase, and only later repurposed into a metric that people used as a proxy for the status of the researchers publishing in those journals. 

I enjoyed hearing about Garfield, who sounds like a delightful and humane polymath, who recognised the value of the information in indexes and found ingenious ways to synthesise it. I can recommend an archive of his work maintained by the University of Pennsylvania.

The later speakers in the webinar focused on novel developments in the use of scientometrics to evaluate research quality. Giovanni Abramo noted how Italian science had been affected by favouritism, because of an exclusive reliance on subjective peer review to evaluate researchers and their institutions. His view is that use of metrics improves research evaluation by making it fairer and more objective. He noted that while metrics may not be an option in some areas of arts and humanities, for disciplines where outputs generally appear in indexed journals, bibliometrics were invaluable, concluding that “bibliometrics are to research assessment what diagnostic imaging is to medicine”- i.e. a key source of objective information. 

Funnily enough, I would have agreed with him 12 years ago, when I suggested a simple bibliometric index (departmental H-index) could achieve very similar results to the complex and time-consuming peer review process adopted in the British Research Excellence Framework. At the time I was writing, I thought that Goodhart’s Law ("When a measure becomes a target, it ceases to be a good measure") wouldn’t apply to a citation-based metric, because citations were not controlled by authors, so it would be difficult to game. 

It turns out I was naïve. The crudest method of gaming is excessive self-citation, but there are also citation rings (you cite my paper and I’ll cite yours). This year Maria Ángeles Oviedo-García, René Aquarius and I described a more sophisticated version, a “review mill”, where a group of Italian medics used positions as peer reviewers to coerce citations to work of the group. We suggested that the change in Italian research evaluation, which had been implemented with the best of intentions, had led to cynical gaming of peer review. One might respond by saying that this activity, though disturbing, affects only a tiny proportion of articles and so would not have a detectable effect. Again, I would have agreed a decade ago. But now, with an explosion in publications that seems driven by publishers who are more focused on income than quality (see Hanson et al, 2024), and extraordinarily lax editorial standards,  that may no longer be true. The key point about review mills is that we saw evidence of their activity because they used generic templates for peer reviews, but these can be detected only for journals that publish open peer review – a tiny minority. The most prolific member of the review mill was a journal editor who had nearly 3000 verified peer reviews listed on Web of Science, only a handful of which we could access. 

So I fear that bibliometrics is more like a diagnostic image that has blurred together data from several patients - there is some valid information in there, but it's distorted by error. 

The final presentation by Valentin Bogorov described the future of Scientometrics, where AI would be harnessed to give much more fine-grained and nuanced information about societal impacts of research. But I felt he was ignoring the elephant of fraud that had lumbered into the bibliometric databases. Review mills are one issue for the validity of citation data, but paper mills are a much bigger problem. Whereas review mills rely on the self-organisation of dubious research groups to burnish their reputations, many paper mills are run by outside organisations whose sole motivation is financial profit (Parker et al., 2024). They sell authorships and citations for a price that depends on the Impact Factor of the journal - Eugene Garfield would be turning in his grave. They were first detected about 12 years ago  but have multiplied like a virus and are seriously infecting whole bodies of research. Sometimes they are first recognised when a researcher with subject-domain expertise finds anomalous or fraudulent articles when trying to review the field (see, e.g., Aquarius et al, 2025). 

Paper mills thrive in a warm, soggy environment, where corrupt or incompetent editors will wave through articles that contain clear breaches of scientific method, or even are evidently patched together from various plagiarised articles. The hope of publishers is that AI will provide ways to detect fraudulent papers and remove them before they enter the literature, but paper millers have proved skilled at mutating to evade detection. Unfortunately, the very areas where AI and big data seem to hold most promise, such as databases linking genes, proteins, molecules and biomarkers, already are contaminated.  The fear is that the papermillers will themselves increasingly use AI to create ever more plausible articles. 

I am not opposed to bibliometrics or AI in principle, but I find the optimism about its application to research evaluation concerning, especially since no mention was made of the problems that will emerge if the database interrogated by AI is polluted. Any method of evaluation will have costs, benefits and unforeseen consequences. My concern is that if we focus only on the benefits, we could end up with a system that encourages fraudsters and rewards those who are most skilled at gaming the system rather than the best scientists.  

Saturday, 18 January 2025

Tomatoes roaming the fields and canaries in the coalmine: another embarrassing paper for MDPI

 


Many publishers are getting nervous about infiltration by paper mills, who can torpedo a journal's reputation when they succeed in publishing papers that are obvious nonsense. In a recent Open Letter, a group of sleuths drew attention to an example in Scientific Reports, published by Springer Nature.

After the Open Letter was published, the paper that instigated our concern was promptly retracted by the journal, but as far as I can tell, not much else has changed. The point about a paper like this is that it is so blatantly bad that it cannot have been through any kind of serious editorial scrutiny or peer review. It acts as a canary in the coalmine: if gobbledegook is published in your journal, it's an indicator that you need to look very carefully at your editorial processes, and act immediately to remove editors who let this stuff in. Sadly, I haven't yet seen much evidence of that happening at Scientific Reports.

This post, however, concerns another publisher, MDPI, who have regularly featured on my blog, and not in a good way. Last month, I commented on the strange state of affairs whereby Finland had downgraded its classification of 187 MDPI journals because of evidence of "minimum time spend for editorial work and quality assessment", at the same time that German universities had secured a national publishing agreement with MDPI. The story I have to tell here may confirm Finland's judgement, and give Germany pause for thought. It concerns this article: Abbas, R., Amran, G. A., Hussain, I., & Ma, S. (2022). A Soft Computing View for the Scientific Categorization of Vegetable Supply Chain Issues. Logistics, 6(3), https://doi.org/10.3390/logistics6030039 

As with the Springer Nature example, the first indication of problems came via the Problematic Paper Screener, the excellent system that checks articles for various red flags, including "tortured phrases". These provide an indicator that a paper has probably been plagiarised but then passed through a process that substitutes synonyms for main words, with the aim of evading plagiarism detection software. So, as noted on PubPeer, in this case we have "fluffy logic" for fuzzy logic, and "unaided ML" for unsupervised machine learning. 

However, example sentences in which tortured phrases were embedded indicated a deeper problem. Most of the text is incomprehensible, and things start to get seriously weird when the authors get on to tomatoes. We are told: 

.... the third creation framework considered for the creation phase is tomatoes. This creation framework is devoted to developing homegrown creatures brought up in rural settings to create vegetables. This can bring domesticated tomatoes likewise to broad or serious frameworks. Broad frameworks include creatures wandering meadows (ordinarily under the oversight of a herder). Differently, serious tomatoes are situated in shut foundations and are outfitted with ICT innovation, which empowers creatures to be observed continuously. Inside these creation frameworks, the most run-of-the-mill issues we run over are meadow observing [75], creature government assistance [76], creature conduct following [77], and tomato creation forecast and enhancement [78,79], as displayed in Figure 3. 

According to a VSC point of view, the formal meanings of these issues are recorded beneath. 

• Field checking: This issue is connected with the exact recognizable proof of meadow inventories to separate between the most reasonable sorts for tomatoes purposes. 

• Tomato government assistance: This is centered around the example arrangement of the dehydration way of behaving in brushing creatures for investigations of creature nourishment, development, and well-being. 

• Tomato growth checking: This depends on the utilization of conduct investigations to recognize early indications of medical problems and advance early negotiation. 

A clue to the origin of this material comes from the cited references, which are about pigs and cattle. Anonymous PubPeer commenter Nerita vitiensis found that a substantial part of the text was adapted from a previous work by different authors, but with the topics of "livestock and fish" changed to "tomatoes and cruciferous vegetables". This explains the description of tomatoes as "creatures" under the oversight of a herder. 

The authors of this piece seem seriously out of their depth, as evidenced by the bland comments apparently written by Chat GPT that they provided on PubPeer. 

Now, one very good thing about MDPI is that it generally identifies the academic editor who handled a paper, and it sometimes also makes public the reviewer reports. This should mean that when a major foul-up like this occurs, it should be possible to identify and purge those responsible for accepting the work. 

The academic editors who accepted this article are Xue-Ming Yuan, who is currently soliciting papers for a special issue in the MDPI journal Mathematics, and Anrong Xue.

The MDPI website shows reports from three named reviewers

The first reviewer, Edyta Kardas, was concerned about the use of first-person language, and punctuation, but not apparently about statements about animated tomatoes. She reviewed 8 papers for MDPI journals in 2024.

The second reviewer, Alejandro Vega-Muñoz focused solely on the structure of the article, but apparently did not look at the content. He has edited two special issues for other MDPI journals

 The third reviewer, Francesco Barreca, attempted a synopsis of the article (which I could not understand) and then had just two suggestions:

"The work is well done but I have some remarks:

• The figures should be review, the dimension are variable

 • Moderate English changes are required"

The "moderate English changes" were unspecified. Barreca has a track record of editing a special issue of another MDPI journal.

 
Last week, I contacted Publication Ethics at MDPI to draw their attention to this article, noting the dereliction of duty by reviewers and editors, and suggesting that as well as retracting the paper, they should remove the editors and peer reviewers from their database. They replied to say: 

"We confirm that the Editorial Office is investigating the concerns related to this paper following the guidelines of the Committee on Publication Ethics https://publicationethics.org/ of which we are a member and our policy https://www.mdpi.com/ethics#_bookmark29.

We would like to inform you that this case is a priority for us, and we are actively working to resolve it. We will update you on the outcome of this investigation as soon as possible."

I await developments with interest. It is widely recognised that COPE guidelines are not well-suited for dealing with this kind of situation: they make the default assumption that authors should be consulted to give their perspective when criticisms are raised - a reasonable assumption in many cases, but not when there is such blatant evidence of fakery.

The most serious case of infestation of a publisher by nonsense occurred in 2022-3, when the publisher Hindawi (owned by Wiley) was targeted by paper mills who, among other things, generated numerous papers that I labelled as AI gobbledegook sandwiches. Eventually, the publisher withdrew literally thousands of papers and closed the Hindawi brand, after complaints by shareholders started impacting profits.

Like many of the sleuths who track down paper mills, I have become cynical about the commitment to research integrity that is claimed by many publishers, including MDPI. But I do believe they will act when it is in their interests to do so. As the amount of nonsense and disinformation in the scientific literature increases, I think we'll enter a new phase where trustworthiness of journal contents will start to have much higher value. If you want to be taken seriously as a peer-reviewed journal, you just cannot continue to pump out articles accompanied by superficial verbiage from "peer reviewers" that makes no real contact with the subject matter. Publishers will need to act now to clean up their editorial boards if they want to stay in business.   


Update 7th February 2025

 

I am pleased to report that I have now heard from MDPI as follows:


Following a thorough investigation of this paper and according to the recommendations of the journal's Editorial Board, we have decided to retract this publication, in line with MDPI’s retraction policy: https://www.mdpi.com/ethics#_bookmark30

Further details regarding this retraction can be found at: https://www.mdpi.com/2305-6290/9/1/20


However, as far as I can tell, no action has been taken against the editors and reviewers who were responsible for accepting the article.  The only thing I notice is that the peer reviewer comments are no longer linked to the main page for the article on the journal website (though they are still available here).  I have written to MDPI Publication Ethics to ask for clarification as to whether any action will be taken to replace the editors, and to remove the peer reviewers from their register.  Unless a robust approach is taken to removing those who admit such nonsense into the journal, readers and potential authors can have no confidence in the integrity of MDPI's editorial processes.

Update 12th February 2025

I have now received a response from MDPI Publication Ethics to my query about editors and peer reviewers.  The plan is to educate them rather than remove them.  


We are happy to provide additional details regarding our approach to this situation. 

In line with the recommendations from COPE for ethical and transparent scholarly publishing (https://publicationethics.org/guidance/guideline/principles-transparency-and-best-practice-scholarly-publishing), we are adopting an educational approach. 
This implies sending notifications to the reviewers and the editor to inform them of this retraction, clarify their responsibilities in the review process, and outline the potential consequences of a superficial review, according to our guidelines for reviewers (https://www.mdpi.com/reviewers#_bookmark11) and information for editors (https://www.mdpi.com/editors). 

Given the significant responsibility that comes with the role of peer reviewers, which directly impacts the credibility of both the journal and the broader scientific literature, we take proactive steps to offer additional guidance and resources beyond our standard guidelines. An example of this initiative can be found on our MDPI Blog: https://blog.mdpi.com/2025/01/22/reviewer-responsibilities/

Moving forward, we will closely monitor the activities of both the editor and reviewers to ensure adherence to the highest standards of academic publishing.

We remain at your disposal for any further information.

Papermillers must be feeling very cheerful right now.

Note: Comments on this blog are moderated, so there is a delay before they appear.  Anonymous or off-topic comments are not accepted.