They are parasites on moral panic the same way “Me too” and wokeness are. Trump saved us from the last two. AI-detection tools are still thriving, they might kill some peoples’ ability to make a living in their profession. Worst of all they stay in the way of huge advantage in research, good writing, translation, produced by AI help.
“Scientists and publishers are testing a new breed of software that promises impressive accuracy at spotting AI-written text.
When Daniel Evanko asked a scientist whether they had used artificial intelligence to write their peer-review report, he didn’t expect a confession. Most researchers don’t reveal AI help, says Evanko, who is the director of journal operations at the American Association for Cancer Research (AACR).
But this time was different. “Wow, you guys are good!” the reviewer wrote back, admitting that he had used a large language model (LLM) after running out of time.
Evanko had a secret weapon. He and the AACR deploy a commercial AI-detection tool called Pangram, because of concerns over the number of peer-review reports submitted to their journals that seem to use AI without disclosing it, contrary to the publisher’s policy.
After years of disappointing results, multiple firms now claim that software can reliably distinguish between AI-written and human-written text. One is Pangram Labs, the New York City-based start-up that makes Pangram. “Detect AI-generated content with 99.98% accuracy,” the firm says on its website.
Scientists and research organizations are among those using the tool to spot AI’s traces. One in eight biomedical articles last year contained some AI-generated text according to Pangram, a study reported in January1. In June, the premier computer-science conference NeurIPS announced that it rejected 18% of submissions after screening them with Pangram. Users of the preprint server arXiv can now check Pangram’s verdict on any article there, by visiting a mirror site called alphaXiv that has installed the tool. And the University of Chicago in Illinois says that it has started using it to vet students’ coursework.
Pangram’s co-founder, Max Spero, says he wants to help everyone spot when text is AI-written. “If it’s taboo to call out that somebody’s using AI to write, then I think we’re going to see a lot more people shirking their jobs and letting AI replace themselves. We’re in a really critical time of setting norms,” he says. Spero has personally called out journalists whom Pangram suggests are using AI and, on one occasion, even flagged the Pope’s social-media posts as AI-written. This July, Pangram was integrated across the popular blogging platform Substack, allowing readers to see whether it deems posts to be AI-written.
Universities are relying on AI-detection software to catch cheating. How well do the programs work?
Pangram isn’t the only firm reporting remarkable results. GPTZero, a competitor also in New York City, says that it has 99% accuracy and provides “the most precise, reliable AI detection results on the market”. Five computer-science conferences and three universities have signed up to use it so far, says the firm’s chief technical officer, Alex Cui, and others are piloting it.
These AI detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, independent analysts say, although no tool can be perfect. And they are “good for screening out places that are pumping out slop”, says Tim Requarth, who studies science communication at New York University’s Langone Health centre in New York City.
But they still sometimes make mistakes, so the tools can be used only as starting points for investigation — and their results are less illuminating for AI-edited writing, in which human and AI text intertwines and there is no clear boundary for problematic use. Nature has tested the tools and interviewed experts to assess how well they do in various situations: where they work, and where they fall short.
How to tell AI writing apart
Until Pangram and others came along, AI-detection software was notoriously unreliable, says Marzena Karpinska, a computer scientist at Simon Fraser University in Burnaby, Canada, who has conducted independent analyses of the tools. Most software analysed statistical characteristics of a text: AI writing tended to display more ‘perplexity’, an estimate of how predictable each word in a sequence is, and less ‘burstiness’, which measures changes in sentence lengths and structures across a text. Some tools would pick up on stylistic quirks in AI writing, such as over-use of the ‘It’s not X, it’s Y’ construction.
But these approaches would too often falsely flag human-authored text as AI-generated. “We didn’t have much faith in the detection,” says Karpinska. Some universities became so fed up with capricious AI detectors that they ended up banning or discouraging staff from using the software.
Then, in early 2024, Spero and Bradley Emi — both computer scientists who had worked at various technology firms since studying together at Stanford University in California — reported a different approach.
They gathered millions of human-written texts and asked LLMs to create ‘mirrors’ of the texts: requesting, for instance, an essay of the same title, length and tone as a human example. Then, they trained machine-learning models to distinguish between the human-written content and the AI mirror, refining the model’s performance by feeding the most challenging cases back in again. Like many such models, the resulting software is a black box. It learns to pick up subtle features of AI writing, often particular to specific LLMs. But these features cannot always be easily described in human terms.
How much of the scientific literature is generated by AI?
In their 2024 study (which has not been peer reviewed)2, Spero and Emi reported a 0.02% false-positive rate: that is, the model rarely labelled human writing as AI. Karpinska and other independent researchers tested the models3. “We were actually quite surprised that it was working really well. Whatever we threw at it, it was really, really good at detection,” she says.
This July, Epoch AI, a research firm in San Francisco, California, reported that Pangram had zero false positives on 495 human-written texts. Pangram’s own latest technical paper reports a 0.0041% false-positive rate on English texts, with similarly low rates in more than 100 languages4. Tests in the study also show that Pangram doesn’t discriminate against writers who are not fluent English-speakers, a problem flagged with earlier software.
By the time Pangram’s early work became public, GPTZero had also switched from measuring perplexity and burstiness to deploying a similar form of machine learning. It, too, performed well in Karpinska’s study and scored zero false positives on Epoch AI’s test. In February, GPTZero marketed a 0.08% false-positive rate in internal tests, although its technical paper5 more cautiously states this as “sub 1%” across all domains of writing.
AI arms race
The tools’ impressive ability to spot solely human-authored work comes with a trade-off. To avoid flagging human text as AI, both firms accept a higher proportion of false negatives — that is, erroneously clearing some AI-generated text as human. Epoch AI found that Pangram and GPTZero flagged nearly all AI-generated passages made using basic prompts.
But when AI was asked to mimic a particular author, some 8% of the resulting passages passed as human.
Karpinska’s study examined a ‘humanizer’ tool — one that rewrites AI-generated text to remove some signs of AI. She found that when detectors were asked to classify human texts versus humanized AI-written text, their false-negative and false-positive rates rose.
But the firms say these tests are already out of date. For instance, Karpinska’s study found that the just-released o1 LLM from OpenAI tripped up an old version of Pangram. Both these tools have now been superseded, and AI-detection firms continually update their models as new versions of LLMs are released.
In the past year, for instance, GPTZero has released 23 updates to its models. Meanwhile, Pangram Labs released its next-generation model, Pangram 4, in July, promising huge advances over its predecessor, including improvements in spotting AI-humanizer signatures. It says its false-negative rate is now only 0.34%, and when AI is asked to imitate styles (as in Epoch’s test), the false-negative rate falls to 2.9%. As with all models, however, performance drops with very short AI passages (of under 50 words).
The upshot is that only the accuracy rates announced by the firms from internal testing can be truly current — but these are not externally verified. “Third-party evaluations will always lag behind the latest detectors,” Requarth says.
The mixed-AI challenge
Both Pangram and GPTZero make money by selling subscriptions for regular or heavy use, but allow some limited free checks. Both have also launched browser tools that check for AI-written text on social media, other web pages and Google docs.
As more writers have experienced being flagged by the software, it’s become clear that the biggest challenge lies in how to approach cases of AI-assisted writing, which is becoming increasingly common. Writers who use AI told Nature that it would be useful to draw a line between using AI to polish or edit human-authored drafts, which they saw as mostly acceptable, and generating a draft with AI from scratch then editing or humanizing it.
Both Pangram and GPTZero break down texts into parts and try to show the degree to which a document is lightly polished with AI, heavily AI-assisted or entirely AI-generated (see ‘How two AI-detectors present their scores’). But in practice, the tools don’t always satisfyingly distinguish between these cases.
For instance, in June, Elena Vicario, director of research integrity at the publisher Frontiers, headquartered in Lausanne, Switzerland, wrote a post for The Scholarly Kitchen, a website that posts views on scientific publishing, arguing that AI can be used to support peer review. When Nature ran this article through Pangram in early July, it came up as 100% AI; after Pangram 4 was released, this changed to 96% AI.
But Vicario says her first draft and the ideas that went into it were “entirely human-created” and that she used AI as a tool to polish the text “which is simply best practice”. Editors at The Scholarly Kitchen add that the article was further edited there, so it couldn’t be wholly AI.
If Pangram labels a work as 100% AI, this doesn’t actually mean every word was AI-generated, Spero says. The tool divides a text into segments and judges whether each segment is probably AI-generated, human-written or ‘mixed’. Then, Pangram assigns an overall score on the basis of the proportion of segments flagged. 100% AI means only that each segment was judged as probably AI, even if the segment has some human-written content.
This chunking process means that a paragraph judged ‘human’ in isolation could switch to AI when combined with other text in a segment, Spero adds — explaining why some writers find that sentences extracted from an article can score differently than when judged in the article as a whole.
“If an editor or author used an AI program to smooth out a sentence or clarify the argument during the editing process, we do not view this as a problem,” Vicario says.
Pangram 4, the software launched this July, has changed evaluations in part because it divides text into more fine-grained sections. Whereas the older version analysed segments of some 200–300 words — meaning that a 1,000-word post might end up being divided into only three sections — the new version can analyse chunks as small as 30–40 words.
Spero says that Pangram 4 better distinguishes between a light AI polish — which usually retains a human-authored label — and heavy AI edits. “It takes major AI input like fully-generated sentences or major rewrites to trigger Pangram 4,” he says.
GPTZero analyses texts slightly differently. Like Pangram, it divides them into segments, but at the end, it gives a score that represents its confidence about the whole text, rather than a breakdown of the results of each segment5. Paid users can see which sentences contributed most to GPTZero’s assessment of a work as probably human, AI or mixed. The tool produces 100% confidence that Vicario’s post was AI, described as meaning the software is “highly confident this text is AI-generated”.
After Nature looked into the case, editors at The Scholarly Kitchen noted on the post that AI had been used “as an editing tool”.
Further questions about Pangram’s evaluations arose when, during the reporting of this article, Spero sent over an examination of articles on Nature’s website. Some articles in a sample from the websites of Nature India and Nature Africa came up as mixed AI–human, and in some cases, 100% AI, he noted, using Pangram 3.3’s assessment.
An examination by the sites’ editorial teams showed that some of the articles flagged were AI-assisted translations of another article or short research highlights that were explicitly written with the aid of AI and had been human-edited. But some were articles written by contributors. They said that they had used AI only to help transcribe or translate interviews, organize notes and edit their drafts. All of these articles went through further rounds of human editing.
In many cases, Pangram 4 judged these articles as merely mixed human–AI, whereas the older version had called them wholly AI. But the newer software still labelled two news reports as 100% AI that contained information derived from original reporting, including interviews. Both sites permit AI use with human oversight, in line with guidance for publications in the Nature Portfolio that says AI should be declared if used for extensive copy-editing or writing, but doesn’t need to be declared for polishing or refining language, although transparency is encouraged. After examining the Pangram results, site editors added notes to some articles to flag the AI assistance.
Overall, as with The Scholarly Kitchen case, the software correctly spotted AI involvement. But Pangram and GPTZero sometimes disagreed on how they rated the articles and on which segments had the strongest indications of being human or AI. And the examples suggest that articles flagged as 100% AI or near 100% can include text written and edited by humans.
Spero says that although Pangram is good at spotting cases in which wholly AI-written paragraphs are inserted between stretches of human-authored text, “there’s a lot of room for improvement” in judging real-world cases in which AI and human writing mingles in a more homogeneous way. For instance, he says, small edits to such documents can sometimes shift Pangram’s score by a large amount — a problem known as ‘jitter’. Pangram’s technical study also notes that when the firm tested using consumer AI to “substantially modify” human-written student essays, the model still labelled the results fully human 41% of the time.
Requarth says that he doesn’t put a lot of faith in exact percentages given by detectors for cases in which AI assistance or editing is involved, and particularly not down to the level of arguing that individual sentences or short chunks of text were AI-written.
The priority for AI detectors should be determining whether a text is fully generated by AI, because “that’s what people are looking for”, says Cui. Determining the extent of AI assistance in a text “is a harder problem”, he says.
A witch hunt?
When Substack readers started to see Pangram’s assessment of posts on the platform, some authors got angry. Sam Illingworth, who studies AI literacy at Edinburgh Napier University, UK, called it a “witch hunt” in a Substack post that Pangram scored as 100% AI. He also ran his text though a humanizer, which switched Pangram’s evaluation to 100% human.
Illingworth says that he makes heavy use of AI tools to edit his drafts. “Absolutely, my work is AI-assisted, but is it 100% AI-generated? No,” he says. The example again shows the potential for confusion with Pangram’s 100%-AI-generated label.
Pangram 4, which was released after Illingworth’s post, rates the original post as 95% AI and 5% human. It assesses the humanized version as 60% AI, “a mix of AI and human-written content” — suggesting that the newer tool does spot traces of AI even in humanized work.
But posts have appeared online suggesting that some humanizing tools can fool Pangram 4, too. “This cat-and-mouse game on both sides is going to continue,” says Nikhil Garg, a computer scientist at Cornell Tech in New York City.
Illingworth’s concerns remain. He says that letting readers check scores on Substack posts might penalize people whose natural writing style is more AI-like (a concern raised in particular by some neurodivergent writers), and punishes those who use AI to improve work if they don’t have English as a first language — a worry echoed by Evanko, who has been testing Pangram on peer reviews.
“The biggest deficiency is when reviewers used LLMs to translate their reviews from non-Indo-European languages into English,” Evanko says. “The result comes back as fully AI generated.” But the AACR is just bringing in Pangram 4, he adds; in early trials, around one-fifth of reviews classed as 100% AI by the older software have now had their AI fraction reduced, with the most extreme reductions for non-Indo-European reviewers.
Spero says that mere translation shouldn’t trigger an AI flag; he thinks that in the cases Evanko mentions, translation and AI editing are occurring simultaneously.
Cui says that the detector’s score shouldn’t be taken as the result in its own right — but just as a signal for further investigation. A pattern of heavy AI use is more revealing than a judgement on a single article, Spero adds.
The inescapable limitation of AI detection, however good it gets, researchers say, is that although software can spot whether AI was involved in a piece, it can’t prove how it was used or judge what’s ethically acceptable.
Readers really want to know whether a piece was meaningfully authored by a human or whether it was low-effort AI slop or spam, but even a perfect AI-spotter can’t prove that case, says Renée DiResta, a researcher at Georgetown University in Washington DC, who studies scams, disinformation campaigns and other examples of online manipulation and abuse. “We end up surfacing what is easiest to detect rather than addressing the deeper underlying concern,” wrote diResta in a blog post about Pangram’s detector on Substack.
Rising distrust
The new tools are also arriving in an online environment awash with poor-quality AI detection. Inferior services continue to circulate, sometimes offering to circumvent AI detectors by humanizing work, and contributing to mistrust of the software. Many critiques of AI detection reference poor-quality tools.
For instance, Mark Carrigan, who works on digital education at the University of Manchester, UK, and has written a book about generative AI for academics, posted online in July that he’d put his old PhD thesis from 2014, written long before LLMs existed, through a tool called TextGuard, which told him that 62% of the file had signs of AI and offered to humanize it.
He concluded that AI detectors couldn’t be trusted; others have reported similar experiences with services such as TextGuard. But when Nature tested Carrigan’s thesis and other texts, GPTZero and Pangram always correctly identified the older content as human-written. Confusingly, TextGuard’s AI score report includes a graphic stating that it is ‘double-checked’ with GPTZero; Cui says GPTZero has no partnership or association with TextGuard. A spokesperson for TextGuard’s creator, the Hong Kong-based firm Level Media Limited, told Nature that this doesn’t mean the firm doesn’t verify results using GPTZero, but declined to clarify further. In general, the spokesperson wrote, TextGuard bases its assessment on similarity to stylistic patterns and sentence structures that are frequently found in AI writing, and false positives can occur.
Pangram’s Spero points out that GPTZero itself was acquired in June by the AI-productivity firm Superhuman, formerly known as Grammarly, which offers AI writing assistance, including tools to humanize AI text. Pangram is the leading independent firm that is not yoked to a humanizer operation, he says.
GPTZero’s Cui responds that the tools are separate and not in conflict; GPTZero even detects whether Superhuman’s products were used. For his part, he argues that Pangram is overconfident in its false-positive rates and that it is spuriously precise in the way that it declares that a particular percentage of content is AI or human. Spero disagrees.
As researchers increasingly use AI to write or edit content and shy away from disclosing it, it’s likely that more publishers will adopt AI screening. Already, the manuscript-screening service Proofig AI in Rehovot, Israel, and the research-integrity firm Clearskies in London are offering Pangram to their users. Of research publishers that replied to Nature’s queries, Science journals say they recently incorporated iThenticate’s AI text-detection feature, and MDPI says that it has built an in-house AI detector named Binoculars. Springer Nature says it is exploring in-house and third-party AI-detection tools. (Nature’s news team is editorially independent of its publisher, Springer Nature.)
In another development, this month, US firm Anthropic said it would introduce watermarking into its Claude AI models to comply with the European Union’s AI Act — although it’s unclear how easily editing could remove the marks, or whether writers will avoid Claude as a result.
Ultimately, it’s likely that if watermarking and AI-detection tools become widespread, writers, including scientists, will need to start more transparently disclosing how they use AI, and to find ways to prove or document their process, says Garg. And that might help everyone find answers to the underlying issue: what counts as ethically acceptable AI use?” [1]
1. AI-detection tools have made huge leaps forward — how good are they? Nature 656, 808-811 (2026) By Miryam Naddaf & Richard Van Noorden
Komentarų nėra:
Rašyti komentarą