Sekėjai

Ieškoti šiame dienoraštyje

2026 m. spalio 2 d., penktadienis

About to give your first lecture? Read this first


“Ten researchers and faculty members share what surprised them when they began teaching — and what they wish they had known before walking into the room.

 

Ľubomír Tomáška still remembers the first university lecture he gave more than 30 years ago. He was a first-year PhD student when a senior teacher asked him to cover an introductory genetics class. The topic was DNA replication: a seemingly “easy ride” for Tomáška, who was doing his doctoral research on a similar topic. He read a chapter from the textbook and prepared a thick stack of transparent slides to project in front of the class — “this was pre-PowerPoint” — and walked in feeling set. “I was ready to go, but in ten minutes, I noticed that I had lost the attention of about 90 per cent of the people in the lecture hall,” recalls Tomáška, now a geneticist at Comenius University in Bratislava. “I felt like a cyclist climbing a steep hill without water and energy.”

 

Tomáška’s experience is familiar to many early-career academics. Across academia, especially during doctoral programmes, the focus is often on fostering research skills, rather than teaching skills. “PhDs teach you to be scientists and researchers, not teachers,” says Olivia Bizimungu, a PhD candidate in neuroscience at McGill University in Montreal, Canada.

 

Despite this, some graduate programmes require students to take on teaching responsibilities, even if they have limited experience doing so. Tomáška, now a sagacious university lecturer, says he often has PhD students enter his programme and immediately begin teaching. But even if you’re very familiar with a subject, it is not all smooth sailing. “Knowing the subject for university teaching is not enough to give a good lecture or seminar,” he adds.

 

Here, ten PhD students, postdocs and faculty members share their advice for anyone preparing to teach at university for the first time.

Build on existing materials

 

Before you embark on your first lesson, you should prepare course materials, such as lecture slides and handouts. Luckily, you don’t have to reinvent the wheel. There is usually an existing course that you can — and should — work off. “Don’t start from zero,” says Aeriel Leonard, a materials scientist the Ohio State University in Columbus. If it’s appropriate, then you might consider using old homework resources or assignments.

 

Tomáška also suggests that researchers see whether it is possible to undertake formal teaching courses at their universities.

 

There’s also another great resource — academics themselves. Like Tomáška, when Melanie Ortiz Alvarez De La Campa — who earned her PhD in pathology at Brown University in Providence, Rhode Island, this May — first taught undergraduate courses, she assumed that having subject knowledge alone would be enough to teach well. Realizing that students process information at different speeds, she sought advice from experienced colleagues. “I learnt a lot from just talking to actual professors,” says Ortiz Alvarez De La Campa. “If you have the opportunity, sit in on some courses where you’re going to teach and see which professors resonate with you in terms of what they’re doing, and talk to them about it.”

Prepare the content and the classroom

 

The students aren’t the only ones learning. Once you have your material, you should also take time to learn it by heart. “Re-read and revise your teaching materials and practise delivering your lectures,” says Nirodha Weeraratne, a postdoc at Charles Sturt University in Wagga Wagga, Australia. “Guess what questions will be asked during discussions, and be ready with plausible answers.”

 

Just as you would do for a conference presentation, you will want to complete a run-through. Liza Bolton, a statistician at the University of Auckland, New Zealand, advises researchers to visit the classroom before their first lesson. “Log into the lectern, test your slides and check your connections,” she says. “Do you need a USB-C or HDMI dongle? What is the microphone situation?”

Black and white photo of Ľubomír Tomáška preparing for a lecture in his younger years.

 

Some 35 years ago, early on in his career, Ľubomír Tomáška prepares for lecturing.Credit: Boris Bilčík

 

But remember, it does not have to be perfect. “The preparation you do for lesson 1 will look entirely different from lesson 100,” adds Bolton. “Don’t despair if you are over-scripting and over-timing right now.”

Engage students from the beginning

 

First impressions are important. Bizimungu suggests grabbing students’ attention early on, however you see fit. “Even if you start with something that there’s no way that they’ll understand, show them why what you’re about to teach is useful and interesting,” she says. For example, when teaching a motor-neuroscience class, Bizimungu started with a video on state-of-the-art brain–computer interfaces. Throughout the lecture, she connected the advanced technology in the video to concepts that the students had already encountered in class.

 

Marco Mello, an ecologist at the University of São Paulo, Brazil, says that you should tell a good story. “Harness the power of storytelling: a clear arc, a conflict, strong characters, a sense of movement and a reason for students to care about what comes next,” he says. “Know where the story begins, where it is going, and what you want students to remember when they leave.”

 

Preparation is also only part of the challenge: Jawad Ullah realized this when he gave his first lecture. “I had to read the room in real time and understand whether they were engaged, sceptical or quietly connecting the ideas to their own classroom experience,” recalls Ullah, a fourth-year PhD candidate in crop science at Hainan University, China. A short pause or a practical question can make the session feel more open, he suggests.

 

Like Tomáška, you might lose students along the way. To reengage them, Ortiz Alvarez De La Campa recommends trying to make the content more personal. “When a student can relate the material directly to something in their real life, or better yet apply it, it helps capture their attention,” she says. “If, as lecturers, we can help them see the real value of the knowledge at work, then they will deem it worth their time.”

Think like the student

 

While you are preparing the lecture, it can be helpful to put yourself in the students’ shoes. Leonard recommends that anytime you give a homework assignment, a quiz or an exam, you should complete it first to make sure that the students have the required knowledge to finish it. You should also get in the students’ mindset. If you’re excited to be there, it can rub off on them. “I always find it really motivating when I'm looking forward to some experiment I'm going to do during the lecture because it keeps my adrenaline high,” says Tomáška.

Melanie Ortiz Alvarez de la Campa standing on a stage delivering a presentation at Brown University.

 

Melanie Ortiz Alvarez De La Campa gives a presentation at Brown University in Providence, Rhode Island.Credit: Brown University Communications Department

 

Teaching is not just about providing course material: it can also involve imparting practical advice. “Remember yourself years ago? What would you need the most at that time?” said Olesia Platonova, a PhD candidate in cognitive neuroscience at the Pasteur Institute in Paris.

Use interactive and varied activities

 

Although lecturing is a tried-and-tested format, you can always vary what you are doing in class to keep it interesting. Kristina Monakhova, a computer scientist at Cornell University in Ithaca, New York, advises taking time to come up with interesting ‘think–pair–share’ activities for each lecture, during which students form individual ideas, discuss with a classmate (or several) and then share with the group.

 

“These interactive discussions are good for breaking up the monotony of a lecture, can help students better engage with the material, and provide you with feedback on whether students are understanding what you're presenting,” she says. Importantly, discussions give you a quick break from active lecturing, allowing you to collect your thoughts and take a water break.

 

To accommodate different levels of confidence, Ortiz Alvarez De La Campa suggests giving students dual assignment options: basic work to complete in class, and then further material if they want to challenge themselves.

Manage your workload and learn from each class

 

As the academic year goes on, Leonard recommends allotting a certain amount of time each day to work on preparing for your lectures — whether that be practising delivering them or grading assignments. She’s learnt from her own experience: “I remember calling my mentor, in the middle of [my first] semester, saying ‘I can't do it, it is really overwhelming,’” she recalls. If you don’t set time limits, she says, "you will recognize that you are being completely overwhelmed by prep and teaching and will feel like you’re drowning”.

 

Bolton suggests using a “dump document” to record ideas and issues. This can be anything from a typo you haven’t fixed to a vibe that a certain lecture just didn’t work. “It can be a mess, but it will help you capture issues that you can reflect on,” she says. “Sometimes, thinking of it as a handover document can help — whether that handover is to future you or a colleague, imagine you want to give them a better experience than you had.”

Accept that improvement comes through practice

 

It can take some trial and error to decide what works best for you. “You have to find those techniques that suit you,” says Tomáška. The more you teach, the more natural it will feel. “While practice doesn’t make perfect, it does make better,” says Bolton.

 

And although it can be nerve-wracking to teach for the first time, you already have the one thing you need: passion. “Passion is the thing that makes it so that people get more engaged in a subject, and especially if you’re someone from a minority background who didn’t see professors like you, it’s even more important,” says Ortiz Alvarez De La Campa. “We need everyone to teach.”" [1]

 

1.   About to give your first lecture? Read this first. By Hannah Docter-Loeb   Nature 658, 565-566 (2026)

Dirbtinio intelekto (DI) aptikimo įrankiai padarė didžiulę pažangą – tačiau kiek jie žalingi?

 


 

Jie parazituoja moralinės panikos dėka – lygiai taip pat, kaip „#MeToo“ judėjimas ar „woke“ kultūra. Nuo pastarųjų dviejų mus išgelbėjo D. Trumpas. DI aptikimo įrankiai vis dar klesti ir gali atimti iš kai kurių žmonių galimybę užsidirbti pragyvenimui savo profesinėje srityje. Blogiausia tai, kad jie trukdo pasinaudoti didžiuliais pranašumais, kuriuos, atliekant tyrimus, kuriant kokybiškus tekstus ar verčiant tekstus, suteikia DI pagalba.


“Mokslininkai ir leidėjai išbando naujos kartos programinę įrangą, kuri žada įspūdingą tikslumą atpažįstant DI parašytus tekstus.

 

Kai Danielis Evanko paklausė vieno mokslininko, ar šis, rengdamas recenzijos ataskaitą, naudojosi dirbtiniu intelektu, jis nesitikėjo prisipažinimo. Pasak D. Evanko, einančio Amerikos vėžio tyrimų asociacijos (AACR) žurnalų veiklos direktoriaus pareigas, dauguma tyrėjų neatskleidžia, kad naudojosi DI pagalba.

 

Tačiau šįkart viskas klostėsi kitaip. „Oho, jūs tikrai šaunūs!“ Recenzentas atsakė pripažindamas, kad pasinaudojo didžiuoju kalbos modeliu (LLM), nes jam pritrūko laiko.

Evanko turėjo slaptą ginklą. Jis kartu su AACR naudoja komercinį dirbtinio intelekto (DI) aptikimo įrankį „Pangram“; šio žingsnio imtasi susirūpinus dėl to, kad į jų žurnalus teikiamose recenzijose vis dažniau, pažeidžiant leidėjo politiką, atrodo, naudojamasi DI apie tai nepranešus.

Po daugelio metų, kai rezultatai nuvildavo, dabar kelios įmonės teigia, kad jų programinė įranga gali patikimai atskirti DI parašytą tekstą nuo žmogaus sukurto teksto. Viena iš tokių įmonių – Niujorke įsikūręs startuolis „Pangram Labs“, sukūręs „Pangram“. Savo interneto svetainėje įmonė skelbia: „Aptikite DI sukurtą turinį 99,98 proc. tikslumu.“

Šį įrankį DI pėdsakams aptikti naudoja mokslininkai ir tyrimų organizacijos. Remiantis „Pangram“ duomenimis, pateiktais sausio mėnesį paskelbtame tyrime, praėjusiais metais kas aštuntame biomedicinos srities straipsnyje buvo aptikta DI sukurto teksto. Birželio mėnesį prestižinė kompiuterių mokslo konferencija „NeurIPS“ pranešė atmetusi 18 proc. pateiktų darbų po to, kai jie buvo patikrinti naudojant „Pangram“. Išankstinių publikacijų (angl. *preprint*) serverio „arXiv“ naudotojai dabar gali sužinoti „Pangram“ išvadą apie bet kurį ten esantį straipsnį apsilankę „alphaXiv“ – veidrodinėje svetainėje, kurioje įdiegtas šis įrankis. Be to, Čikagos universitetas (Ilinojaus valstija) teigia pradėjęs naudoti šią priemonę studentų rašto darbams tikrinti.

„Pangram“ bendraįkūrėjas Maxas Spero sako norintis padėti visiems atpažinti DI parašytą tekstą. „Jei laikoma tabu viešai įvardyti, kad kas nors tekstui rašyti naudoja DI, manau, pamatysime vis daugiau žmonių, vengiančių savo darbo ir leidžiančių DI juos pakeisti.“ „Gyvename itin svarbiu normų nustatymo laikotarpiu“, – teigia jis. Spero asmeniškai nurodė konkrečius žurnalistus, kurie, „Pangram“ vertinimu, naudojasi dirbtiniu intelektu (DI), o kartą netgi pažymėjo popiežiaus įrašus socialiniuose tinkluose kaip sukurtus DI. Šių metų liepą „Pangram“ buvo integruota į populiarią tinklaraščių platformą „Substack“; tai leidžia skaitytojams matyti, ar sistema laiko įrašus sukurtus DI pagalba.

 

Universitetai, siekdami aptikti sukčiavimą, pasitelkia DI atpažinimo programinę įrangą. Kaip gerai šios programos veikia?

 

„Pangram“ nėra vienintelė įmonė, demonstruojanti įspūdingus rezultatus. „GPTZero“ – konkurentė, taip pat įsikūrusi Niujorke, – teigia pasiekianti 99 proc. tikslumą ir užtikrinanti „tiksliausius bei patikimiausius DI atpažinimo rezultatus rinkoje“. Pasak įmonės technikos direktoriaus Alexo Cui, šią sistemą jau nusprendė naudoti penkios kompiuterių mokslo konferencijos ir trys universitetai, o kiti ją išbando.

Nepriklausomų analitikų teigimu, šie DI detektoriai veikia – jie beveik visada teisingai atpažįsta tik žmonių sukurtą turinį, nors jokia priemonė negali būti tobula. Vis dėlto jie „praverčia atskiriant atvejus, kai masiškai generuojamas prastos kokybės turinys“, – sako Timas Requarthas, tiriantis mokslo komunikaciją Niujorko universiteto „Langone Health“ centre.

Tačiau klaidos vis tiek pasitaiko, todėl šios priemonės gali būti naudojamos tik kaip atspirties taškas tyrimui; be to, jų rezultatai ne tokie informatyvūs vertinant tekstus, kuriuos redagavo DI – kai žmogaus ir DI tekstas susipina ir nėra aiškios ribos, žyminčios probleminį naudojimą. Žurnalas „Nature“ išbandė šias priemones ir pakalbino ekspertus, siekdamas įvertinti jų veikimą įvairiose situacijose: kur jos veikia gerai, o kur – ne.

Kaip atpažinti DI sukurtą tekstą

Iki pasirodant „Pangram“ ir kitoms panašioms sistemoms, DI atpažinimo programinė įranga garsėjo nepatikimumu, teigia Marzena Karpinska, kompiuterių mokslininkė iš Simono Fraserio universiteto Burnabyje (Kanada), atlikusi nepriklausomą šių priemonių analizę. Dauguma programų analizavo statistines teksto charakteristikas: DI kuriamiems tekstams būdingas didesnis „painumo“ (angl. *perplexity*) lygis – rodiklis, parodantis, kiek nuspėjamas yra kiekvienas žodis sekoje, – ir mažesnis „netolygumo“ (angl. *burstiness*) lygis, matuojantis sakinių ilgio pokyčius bei struktūras visame tekste. Kai kurie įrankiai atpažindavo dirbtinio intelekto kuriamiems tekstams būdingus stilistinius ypatumus, pavyzdžiui, pernelyg dažną konstrukcijos „tai ne X, o Y“ vartojimą.

Tačiau taikant šiuos metodus dažnai klaidingai nustatoma, kad žmogaus parašytas tekstas yra sukurtas dirbtinio intelekto. „Mes nelabai pasitikėjome šiuo aptikimo būdu“, – teigia Karpinska. Kai kurie universitetai taip nusivylė nepatikimais dirbtinio intelekto detektoriais, kad galiausiai uždraudė darbuotojams jais naudotis arba atkalbinėjo juos nuo to.

 

Vėliau, 2024 m. pradžioje, Spero ir Bradley Emi – abu kompiuterių mokslininkai, dirbę įvairiose technologijų įmonėse po bendrų studijų Stanfordo universitete Kalifornijoje, – pristatė kitokį metodą.

Jie surinko milijonus žmonių parašytų tekstų ir paprašė didžiųjų kalbos modelių (LLM) sukurti tų tekstų „veidrodines kopijas“: pavyzdžiui, prašydami parašyti tokio paties pavadinimo, ilgio ir tono esė, kokia buvo žmogaus sukurtas pavyzdys. Tada jie apmokė mašininio mokymosi modelius atskirti žmonių parašytą turinį nuo DI sukurtų kopijų, tobulindami modelio veikimą pakartotinai įvedant sudėtingiausius atvejus. Kaip ir daugelis tokių modelių, sukurta programinė įranga veikia kaip „juodoji dėžė“. Ji išmoksta atpažinti subtilius DI rašymo požymius, dažnai būdingus konkretiems didiesiems kalbos modeliams. Tačiau šių požymių ne visada įmanoma lengvai apibūdinti žmonėms suprantamais terminais.

 

Kokia mokslo literatūros dalis yra sukurta dirbtinio intelekto?

 

Savo 2024 m. tyrime (kuris nebuvo recenzuotas)2 Spero ir Emi nurodė 0,02 proc. klaidingai teigiamų rezultatų rodiklį: tai reiškia, kad modelis retai klaidingai priskirdavo žmonių tekstus DI kategorijai. Karpinska ir kiti nepriklausomi tyrėjai išbandė šiuos modelius3. „Iš tiesų buvome gana nustebinti, kad sistema veikė taip gerai. Kad ir ką jai pateikdavome, ji puikiai atpažindavo tekstus“, – teigia ji.

Šių metų liepą San Franciske (Kalifornijoje) įsikūrusi tyrimų bendrovė „Epoch AI“ pranešė, kad „Pangram“ nepadarė nė vienos klaidos (nulis klaidingai teigiamų rezultatų), tikrinant 495 žmonių parašytus tekstus. Naujausiame techniniame pranešime apie „Pangram“ nurodomas 0,0041 proc. klaidingai teigiamų rezultatų rodiklis tikrinant tekstus anglų kalba; panašiai žemi rodikliai fiksuojami ir daugiau nei 100 kitų kalbų4. Tyrimo metu atlikti bandymai taip pat rodo, kad „Pangram“ nediskriminuoja autorių, kuriems anglų kalba nėra gimtoji – tai problema, būdinga ankstesnei programinei įrangai.

Kol „Pangram“ kūrimo darbai dar buvo pradinėje stadijoje, „GPTZero“ taip pat pakeitė veikimo principą: atsisakė „sumišimo“ (angl. *perplexity*) ir „teksto srauto netolygumo“ (angl. *burstiness*) matavimo metodų ir pradėjo naudoti panašų mašininio mokymosi būdą. Šis įrankis taip pat gerai pasirodė Karpinskos tyrime ir „Epoch AI“ atliktame teste nepadarė nė vienos klaidos. Vasario mėnesį „GPTZero“ skelbė, kad atliekant vidinius bandymus klaidingai teigiamų rezultatų (kai žmogaus tekstas klaidingai palaikomas dirbtinio intelekto tekstu) dalis siekė 0,08 proc., nors techninėje ataskaitoje atsargiau nurodoma, kad visose rašymo srityse šis rodiklis yra „mažesnis nei 1 proc.“.

DI technologijų lenktynės

Įspūdingas šių įrankių gebėjimas atpažinti tik žmonių sukurtus tekstus turi ir trūkumų. Siekdamos išvengti atvejų, kai žmogaus tekstas klaidingai priskiriamas DI, abi įmonės toleruoja didesnę klaidingai neigiamų rezultatų dalį – t. y. kai DI sukurtas tekstas klaidingai palaikomas žmogaus kūriniu. „Epoch AI“ nustatė, kad „Pangram“ ir „GPTZero“ atpažino beveik visas DI sukurtas teksto ištraukas, parengtas naudojant paprastas užklausas.

 

Tačiau kai DI buvo paprašyta imituoti konkretų autorių, apie 8 proc. gautų ištraukų buvo palaikytos žmonių parašytais tekstais.

Karpinskos tyrime nagrinėtas „humanizavimo“ įrankis – priemonė, perrašanti DI sukurtą tekstą taip, kad jame sumažėtų DI būdingų požymių. Tyrėja nustatė, kad kai detektorių buvo prašoma atskirti žmonių tekstus nuo „humanizuotų“ DI tekstų, klaidingai neigiamų ir klaidingai teigiamų rezultatų rodikliai išaugo.

Vis dėlto įmonės teigia, kad šie bandymų duomenys jau pasenę. Pavyzdžiui, Karpinskos tyrimas parodė, kad ką tik „OpenAI“ išleistas didysis kalbos modelis (LLM) „o1“ suklaidino senąją „Pangram“ versiją. Abu šie įrankiai jau pakeisti naujesniais, o DI aptikimo įmonės nuolat atnaujina savo modelius, pasirodžius naujoms LLM versijoms.

Pavyzdžiui, per pastaruosius metus „GPTZero“ savo modeliams išleido 23 atnaujinimus. Tuo tarpu „Pangram Labs“ liepos mėnesį pristatė naujos kartos modelį „Pangram 4“, žadėdama didelę pažangą lyginant su pirmtaku, įskaitant geresnį gebėjimą atpažinti DI tekstų „humanizavimo“ požymius. Įmonė teigia, kad dabar klaidingai neigiamų rezultatų dalis siekia vos 0,34 proc., o kai DI prašoma imituoti tam tikrą stilių (kaip „Epoch“ bandyme), šis rodiklis sudaro 2,9 proc. Visgi, kaip ir kitų modelių atveju, veikimo tikslumas sumažėja apdorojant labai trumpas DI sukurtas ištraukas (trumpesnes nei 50 žodžių).

Taigi, tiksliausius ir aktualiausius duomenis apie veikimo tikslumą pateikia pačios įmonės, remdamosi vidiniais bandymais, tačiau šie duomenys nėra nepriklausomai patikrinti. „Trečiųjų šalių vertinimai visada atsiliks nuo naujausių detektorių“, – sako Requarthas.

 

Mišraus dirbtinio intelekto iššūkis

Tiek „Pangram“, tiek „GPTZero“ uždirba pinigus parduodamos prenumeratas reguliariam ar intensyviam naudojimui, tačiau leidžia atlikti tam tikrus ribotus nemokamus patikrinimus. Abi taip pat pristatė naršyklės įrankius, kurie tikrina dirbtinio intelekto parašytą tekstą socialinėje žiniasklaidoje, kituose tinklalapiuose ir „Google“ dokumentuose.

 

Kadangi vis daugiau rašytojų susiduria su programinės įrangos žymėjimu, tapo aišku, kad didžiausias iššūkis yra tai, kaip spręsti dirbtinio intelekto padedamo rašymo atvejus, kurie tampa vis dažnesni. Rašytojai, kurie naudoja dirbtinį intelektą, žurnalui „Nature“ sakė, kad būtų naudinga nubrėžti ribą tarp dirbtinio intelekto naudojimo žmonių parašytų juodraščių redagavimui ar šlifavimui, kurį jie laikė dažniausiai priimtinu, ir juodraščio generavimo su dirbtiniu intelektu nuo nulio, o vėliau jo redagavimo ar humanizavimo.

 

Tiek „Pangram“, tiek „GPTZero“ suskirsto tekstus į dalis ir bando parodyti, kiek dokumentas yra lengvai šlifuotas dirbtinio intelekto, stipriai dirbtinio intelekto padedamas ar visiškai sukurtas dirbtinio intelekto (žr. „Kaip du dirbtinio intelekto detektoriai pateikia savo balus“). Tačiau praktiškai įrankiai ne visada patenkinamai atskiria šiuos atvejus.

 

Pavyzdžiui, birželio mėn. Elena Vicario, „Social Health and Health“ direktorė, atsakinga už  už tyrimų sąžiningumą atsakinga „Frontiers“ (leidyklos, kurios būstinė yra Lozanoje, Šveicarijoje) darbuotoja parašė straipsnį svetainei „The Scholarly Kitchen“, kurioje skelbiamos nuomonės apie mokslinę leidybą; jame ji teigė, kad dirbtinis intelektas (DI) gali būti naudojamas padedant atlikti recenzavimo procesą. Kai liepos pradžioje „Nature“ patikrino šį straipsnį naudodama „Pangram“ įrankį, šis parodė, kad tekstas 100 proc. sukurtas DI; išleidus „Pangram 4“ versiją, šis rodiklis pakito iki 96 proc.

Tačiau Vicario teigia, kad pirminis juodraštis ir jame išdėstytos idėjos buvo „visiškai sukurti žmogaus“, o DI ji pasitelkė tik kaip įrankį tekstui patobulinti – „o tai tiesiog atitinka geriausią praktiką“. „The Scholarly Kitchen“ redaktoriai priduria, kad straipsnis buvo papildomai redaguojamas ir jų svetainėje, tad jis negalėjo būti visiškai sukurtas DI.

Pasak Spero, jei „Pangram“ įvertina tekstą kaip 100 proc. sukurtą DI, tai nebūtinai reiškia, kad kiekvieną žodį sugeneravo DI. Įrankis suskirsto tekstą į segmentus ir vertina, ar kiekvienas segmentas greičiausiai sukurtas DI, parašytas žmogaus, ar yra „mišrus“. Tada „Pangram“ pateikia bendrą įvertinimą, pagrįstą pažymėtų segmentų dalimi. 100 proc. DI rodiklis reiškia tik tai, kad kiekvienas segmentas buvo įvertintas kaip greičiausiai sukurtas DI, net jei jame yra žmogaus parašyto turinio.

Dėl tokio skirstymo į dalis pastraipa, kuri atskirai vertinant būtų priskirta žmogaus kūrybai, sujungus ją su kitu tekstu segmente gali būti priskirta DI, priduria Spero. Tai paaiškina, kodėl kai kurie autoriai pastebi, kad atskirai paimti sakiniai iš straipsnio gali gauti kitokį įvertinimą nei vertinant visą straipsnį.

„Jei redaktorius ar autorius redagavimo metu pasitelkė DI programą sakiniui patobulinti ar argumentui patikslinti, mes nelaikome to problema“, – sako Vicario.

Šių metų liepą pristatyta programinė įranga „Pangram 4“ pakeitė vertinimo rezultatus iš dalies dėl to, kad tekstą skirsto į smulkesnes dalis. Ankstesnė versija analizavo maždaug 200–300 žodžių segmentus (tai reiškė, kad 1 000 žodžių straipsnis galėjo būti suskirstytas tik į tris dalis), o naujoji versija gali analizuoti vos 30–40 žodžių apimties atkarpas. Spero teigia, kad „Pangram 4“ geriau atskiria nedidelį dirbtinio intelekto (DI) atliktą teksto šlifavimą – kai paprastai išlaikoma nuoroda į žmogų kaip autorių – nuo ​​esminių DI atliktų pakeitimų. „Kad suveiktų „Pangram 4“, reikalingas didelis DI indėlis, pavyzdžiui, visiškai sugeneruoti sakiniai arba iš esmės perrašytas tekstas“, – sako jis.

„GPTZero“ tekstus analizuoja kiek kitaip. Kaip ir „Pangram“, ši programa suskirsto tekstus į segmentus, tačiau galutiniame etape pateikia įvertinimą, rodantį pasitikėjimo visą tekstą apimančia išvada laipsnį, o ne atskirų segmentų analizės rezultatus. Mokami vartotojai gali matyti, kurie sakiniai labiausiai lėmė „GPTZero“ sprendimą priskirti tekstą žmogaus, DI ar mišriai kategorijai. Įrankis parodė 100 proc. tikrumą, kad Vicario įrašas sukurtas naudojant DI; tai reiškia, kad programinė įranga yra „labai tikra, jog šis tekstas sugeneruotas DI“.

 

 

„Nature“ išnagrinėjus šį atvejį, „The Scholarly Kitchen“ redaktoriai prie įrašo pridėjo pastabą, kad DI buvo naudojamas „kaip redagavimo įrankis“.

Papildomų klausimų dėl „Pangram“ vertinimų kilo, kai rengiant šį straipsnį Spero pateikė „Nature“ svetainėje skelbtų straipsnių analizę. Remiantis „Pangram 3.3“ vertinimu, kai kurie straipsniai iš „Nature India“ ir „Nature Africa“ svetainių buvo priskirti mišriai (DI ir žmogaus) kategorijai, o kai kuriais atvejais – visiškai (100 proc.) DI sugeneruotų tekstų kategorijai.

Šių svetainių redakcijų atliktas patikrinimas parodė, kad dalis pažymėtų straipsnių buvo kitų straipsnių vertimai, atlikti padedant DI, arba trumpos tyrimų apžvalgos, kurios buvo specialiai rengiamos pasitelkiant DI ir vėliau redaguojamos žmonių. Tačiau būta ir straipsnių, kuriuos parašė išoriniai autoriai. Jie teigė, kad DI naudojo tik padėdami transkribuoti ar versti interviu, tvarkyti užrašus ir redaguoti juodraščius. Visi šie straipsniai vėliau buvo papildomai redaguojami žmonių.

Daugeliu atvejų „Pangram 4“ šiuos straipsnius įvertino kaip mišrius (žmogaus ir DI), nors ankstesnė versija juos priskyrė visiškai DI sugeneruotų tekstų kategorijai. Vis dėlto naujesnė programinė įranga vis tiek kaip 100 proc. DI sukurtus pažymėjo du naujienų pranešimus, kuriuose buvo pateikta informacija, gauta atlikus originalų tyrimą ar interviu. Abiejose svetainėse leidžiama naudoti dirbtinį intelektą (DI) prižiūrint žmogui – tai atitinka „Nature Portfolio“ leidinių gaires, kuriose nurodoma, kad apie DI naudojimą būtina pranešti, jei jis pasitelkiamas išsamiam teksto redagavimui ar rašymui, tačiau to nereikia darant nedidelius kalbinius pataisymus ar tobulinant tekstą (nors skatinamas skaidrumas). Išanalizavę „Pangram“ rezultatus, svetainių redaktoriai prie kai kurių straipsnių pridėjo pastabas, nurodančias, kad buvo pasitelktas DI.

Apskritai, kaip ir „The Scholarly Kitchen“ atveju, programinė įranga teisingai nustatė DI dalyvavimą. Vis dėlto „Pangram“ ir „GPTZero“ kartais nesutardavo vertindami straipsnius bei nustatydami, kuriose teksto dalyse labiausiai jaučiamas žmogaus ar DI „braižas“. Be to, pavyzdžiai rodo, kad straipsniuose, kuriuos įrankiai įvertino kaip 100 proc. (ar beveik 100 proc.) sukurtus DI, gali būti ir žmonių parašyto bei redaguoto teksto.

Spero teigia, kad nors „Pangram“ gerai atpažįsta atvejus, kai tarp žmogaus parašyto teksto įterpiamos visiškai DI sukurtos pastraipos, vertinant realias situacijas, kuriose DI ir žmogaus tekstas susipina tolygiau, „dar yra daug erdvės tobulėjimui“. Pavyzdžiui, pasak jo, atlikus nedidelius tokių dokumentų pakeitimus, „Pangram“ pateikiamas įvertinimas kartais gali smarkiai pasikeisti kyla problema, vadinama „virpėjimu“ (angl. *jitter*). „Pangram“ techninėje ataskaitoje taip pat pažymima, kad kai įmonė išbandė vartotojams skirtus dirbtinio intelekto (DI) įrankius, norėdama „iš esmės pakeisti“ žmonių parašytus studentų rašinius, modelis vis tiek 41 proc. atvejų rezultatus įvertino kaip visiškai sukurtus žmogaus.

Requarthas teigia ne itin pasitikintis tiksliais procentiniais rodikliais, kuriuos pateikia detektoriai tais atvejais, kai naudojama DI pagalba ar redagavimas – ypač kai bandoma nustatyti, ar atskirus sakinius bei trumpas teksto atkarpas parašė DI.

DI detektorių prioritetas turėtų būti nustatyti, ar tekstas visiškai sukurtas DI, nes „būtent to žmonės ir ieško“, – sako Cui. Nustatyti DI pagalbos mastą tekste yra „sudėtingesnis uždavinys“, – priduria jis.

Raganų medžioklė?

Kai „Substack“ skaitytojai pradėjo matyti „Pangram“ atliekamą platformoje skelbiamų įrašų vertinimą, kai kurie autoriai pasipiktino. Samas Illingworthas, Edinburgo „Napier“ universitete (JK) tiriantis DI raštingumą, pavadino tai „raganų medžiokle“ įraše, kurį „Pangram“ įvertino kaip 100 proc. sukurtą DI. Jis taip pat perleido savo tekstą per „humanizavimo“ įrankį, po kurio „Pangram“ vertinimas pasikeitė į 100 proc. sukurtą žmogaus.

Illingworthas teigia dažnai naudojantis DI įrankius juodraščiams redaguoti. „Be abejo, mano darbe naudojama DI pagalba, tačiau ar jis 100 proc. sukurtas DI? Ne“, – sako jis. Šis pavyzdys dar kartą parodo, kokių nesusipratimų gali kilti dėl „Pangram“ naudojamos „100 proc. DI sukurtas“ žymos.

„Pangram 4“ versija, išleista jau po Illingwortho įrašo, pradinį įrašą įvertina kaip 95 proc. DI ir 5 proc. žmogaus kūrinį. „Humanizuotą“ versiją ji įvertina kaip 60 proc. DI – tai „DI ir žmogaus parašyto turinio derinys“, rodantis, kad naujesnis įrankis pastebi DI pėdsakų net ir „humanizuotame“ tekste.

 

 

Vis dėlto internete pasirodė pranešimų, teigiančių, kad kai kurie „humanizavimo“ įrankiai gali apgauti ir „Pangram 4“. „Šis abipusis „katės ir pelės“ žaidimas tęsis“, – teigia Niujorko „Cornell Tech“ kompiuterių mokslininkas Nikhilas Gargas.

Illingworthas vis dar turi nuogąstavimų. Jis teigia, kad leidimas skaitytojams matyti „Substack“ įrašų vertinimo balus gali pakenkti žmonėms, kurių natūralus rašymo stilius primena dirbtinio intelekto (DI) kuriamą tekstą (tai ypač aktualu kai kuriems neurodivergentiškiems autoriams), taip pat nubausti tuos, kurie pasitelkia DI tekstui pagerinti, nors anglų kalba jiems nėra gimtoji. Šiai nuomonei pritaria ir Evanko, bandęs „Pangram“ įrankį recenzijų vertinimui.

„Didžiausia problema kyla tada, kai recenzentai naudoja didžiuosius kalbos modelius (LLM) savo recenzijoms iš ne indoeuropiečių kalbų į anglų kalbą išversti, – sako Evanko. – Rezultatas atrodo taip, lyg būtų visiškai sukurtas DI.“ Visgi jis priduria, kad AACR pradeda naudoti „Pangram 4“ versiją; pradinių bandymų metu maždaug penktadalio recenzijų, kurias senesnė programinė įranga priskyrė prie 100 proc. sukurtų DI, DI dalis sumažėjo – ryškiausias sumažėjimas pastebėtas recenzentų, rašančių ne indoeuropiečių kalbomis, atveju.

Spero teigia, kad vien vertimas neturėtų lemti DI naudojimo žymos atsiradimo; jo manymu, Evanko minimais atvejais vertimas ir redagavimas pasitelkiant DI vyksta vienu metu.

Cui pažymi, kad detektoriaus pateikiamo balo nereikėtų laikyti galutine išvada – tai tik signalas, rodantis, kad reikalingas išsamesnis tyrimas. Pasak Spero, nuolatinis intensyvus DI naudojimas yra informatyvesnis rodiklis nei vieno konkretaus straipsnio vertinimas.

Tyrėjų teigimu, neišvengiamas DI aptikimo priemonių trūkumas – kad ir kokios pažangios jos būtų – yra tas, jog programinė įranga gali nustatyti, ar kuriant tekstą dalyvavo DI, tačiau negali įrodyti, kaip jis buvo naudojamas, ar įvertinti, kas yra etiškai priimtina.

Skaitytojai labai nori žinoti, ar tekstą iš esmės parašė žmogus, ar tai – mažai pastangų reikalaujantis, prastos kokybės DI sugeneruotas turinys ar brukalas, tačiau net ir tobulas DI atpažinimo įrankis negali to įrodyti, teigia Renée DiResta, Džordžtauno universiteto (Vašingtonas) tyrėja, nagrinėjanti sukčiavimo atvejus, dezinformacijos kampanijas bei kitus manipuliavimo ir piktnaudžiavimo internete pavyzdžius. „Galiausiai mes iškeliame į paviršių tai, ką lengviausia aptikti, užuot sprendę gilesnę problemą“, – rašė diResta tinklaraščio įraše apie „Pangram“ detektorių „Substack“ svetainėje.

 

Didėjantis nepasitikėjimas

Naujieji įrankiai taip pat atkeliauja į internetinę aplinką, kurioje gausu prastos kokybės dirbtinio intelekto aptikimo. Toliau cirkuliuoja prastos kokybės paslaugos, kartais siūlančios apeiti dirbtinio intelekto detektorius, sužmoginant darbą ir prisidedančios prie nepasitikėjimo programine įranga. Daugelyje dirbtinio intelekto aptikimo kritikų minimi prastos kokybės įrankiai.

Pavyzdžiui, Markas Carriganas, dirbantis skaitmeninio švietimo srityje Mančesterio universitete, JK, ir parašęs knygą apie generatyvinį dirbtinį intelektą akademikams, liepą internete paskelbė, kad savo seną daktaro disertaciją, parašytą 2014 m., gerokai prieš atsirandant teisės magistro laipsniams, jis įkėlė per įrankį, vadinamą „TextGuard“. Šis įrankis jam pasakė, kad 62 % failo turi dirbtinio intelekto požymių, ir pasiūlė jį sužmoginti.

 

Jis padarė išvadą, kad dirbtinio intelekto detektoriais negalima pasitikėti; kiti pranešė apie panašią patirtį su tokiomis paslaugomis kaip „TextGuard“. Tačiau kai „Nature“ testavo Carrigan tezę ir kitus tekstus, „GPTZero“ ir „Pangram“ visada teisingai atpažino senesnį turinį kaip žmogaus parašytą. Keista, bet „TextGuard“ dirbtinio intelekto vertinimo ataskaitoje yra grafikas, kuriame teigiama, kad jis „dvigubai patikrintas“ su „GPTZero“; Cui teigia, kad „GPTZero“ neturi jokios partnerystės ar ryšio su „TextGuard“. „TextGuard“ kūrėjo, Honkongo „Level Media Limited“ atstovas leidiniui „Nature“ teigė, kad tai nereiškia, jog įmonė netikrina rezultatų naudodama „GPTZero“, tačiau atsisakė pateikti daugiau paaiškinimų. Pasak atstovo, „TextGuard“ savo vertinimą grindžia panašumu į stilistinius modelius ir sakinių struktūras, būdingus dirbtinio intelekto (DI) kuriamiems tekstams, todėl galimi ir klaidingai teigiami rezultatai.

„Pangram“ atstovas Spero pažymi, kad pačią „GPTZero“ birželio mėnesį įsigijo DI produktyvumo įrankius kurianti bendrovė „Superhuman“ (anksčiau žinoma kaip „Grammarly“), siūlanti pagalbą rašant tekstus, įskaitant įrankius, skirtus DI sukurtam tekstui suteikti „žmogišką“ skambesį. Jo teigimu, „Pangram“ yra pirmaujanti nepriklausoma įmonė, nesusijusi su DI tekstų „humanizavimo“ veikla.

„GPTZero“ atstovas Cui atsako, kad šie įrankiai yra atskiri ir tarpusavyje nekonfliktuoja; „GPTZero“ netgi geba aptikti, ar buvo naudoti „Superhuman“ produktai. Savo ruožtu jis teigia, kad „Pangram“ pernelyg pasitiki savo klaidingai teigiamų rezultatų rodikliais ir nepagrįstai tiksliai nurodo, kokią turinio dalį sukūrė DI, o kokią – žmogus. Spero su tuo nesutinka.

Tyrėjams vis dažniau naudojant DI turiniui kurti ar redaguoti ir vengiant tai atskleisti, tikėtina, kad daugiau leidėjų pradės taikyti DI aptikimo priemones. Jau dabar rankraščių tikrinimo paslauga „Proofig AI“ (Rechovotas, Izraelis) ir tyrimų vientisumą užtikrinanti įmonė „Clearskies“ (Londonas) savo vartotojams siūlo „Pangram“ įrankį. Iš į „Nature“ užklausas atsakiusių mokslo leidėjų, „Science“ žurnalų grupė nurodė neseniai įdiegusi „iThenticate“ DI teksto aptikimo funkciją, o MDPI teigia sukūrusi nuosavą DI detektorių „Binoculars“. „Springer Nature“ praneša svarstanti galimybę naudoti tiek savo, tiek trečiųjų šalių sukurtus DI aptikimo įrankius. („Nature“ naujienų skyrius redakciniu požiūriu yra nepriklausomas nuo leidėjo „Springer Nature“.)

Be to, šį mėnesį JAV bendrovė „Anthropic“ paskelbė, kad į savo DI modelius „Claude“ įdiegs skaitmenines žymes (angl. *watermarking*), siekdama laikytis Europos Sąjungos Dirbtinio intelekto akto reikalavimų, nors kol kas neaišku, kaip lengvai redaguojant būtų galima pašalinti šias žymes ir ar dėl to rašantieji vengs naudotis „Claude“. Galiausiai tikėtina, kad jei vandens ženklų naudojimas ir dirbtinio intelekto (DI) aptikimo įrankiai taps plačiai paplitę, rašytojams, įskaitant mokslininkus, teks skaidriau atskleisti, kaip jie naudoja DI, bei rasti būdų pagrįsti ar dokumentuoti savo darbo procesą, teigia Garg. Tai gali padėti visiems rasti atsakymą į esminį klausimą: kas laikoma etiškai priimtinu DI naudojimu?” [1]

 

1. AI-detection tools have made huge leaps forward — how good are they? Nature 656, 808-811 (2026) By Miryam Naddaf & Richard Van Noorden

 

AI-detection tools have made huge leaps forward — how harmful are they?


They are parasites on moral panic the same way “Me too” and wokeness are. Trump saved us from the last two. AI-detection tools are still thriving, they might kill some peoples’ ability to make a living in their profession. Worst of all they stay in the way of huge advantage in research, good writing, translation, produced by AI help.

 

 “Scientists and publishers are testing a new breed of software that promises impressive accuracy at spotting AI-written text.

 

When Daniel Evanko asked a scientist whether they had used artificial intelligence to write their peer-review report, he didn’t expect a confession. Most researchers don’t reveal AI help, says Evanko, who is the director of journal operations at the American Association for Cancer Research (AACR).

 

 

But this time was different. “Wow, you guys are good!” the reviewer wrote back, admitting that he had used a large language model (LLM) after running out of time.

Evanko had a secret weapon. He and the AACR deploy a commercial AI-detection tool called Pangram, because of concerns over the number of peer-review reports submitted to their journals that seem to use AI without disclosing it, contrary to the publisher’s policy.

After years of disappointing results, multiple firms now claim that software can reliably distinguish between AI-written and human-written text. One is Pangram Labs, the New York City-based start-up that makes Pangram. “Detect AI-generated content with 99.98% accuracy,” the firm says on its website.

Scientists and research organizations are among those using the tool to spot AI’s traces. One in eight biomedical articles last year contained some AI-generated text according to Pangram, a study reported in January1. In June, the premier computer-science conference NeurIPS announced that it rejected 18% of submissions after screening them with Pangram. Users of the preprint server arXiv can now check Pangram’s verdict on any article there, by visiting a mirror site called alphaXiv that has installed the tool. And the University of Chicago in Illinois says that it has started using it to vet students’ coursework.

Pangram’s co-founder, Max Spero, says he wants to help everyone spot when text is AI-written. “If it’s taboo to call out that somebody’s using AI to write, then I think we’re going to see a lot more people shirking their jobs and letting AI replace themselves. We’re in a really critical time of setting norms,” he says. Spero has personally called out journalists whom Pangram suggests are using AI and, on one occasion, even flagged the Pope’s social-media posts as AI-written. This July, Pangram was integrated across the popular blogging platform Substack, allowing readers to see whether it deems posts to be AI-written.

 

Universities are relying on AI-detection software to catch cheating. How well do the programs work?

 

Pangram isn’t the only firm reporting remarkable results. GPTZero, a competitor also in New York City, says that it has 99% accuracy and provides “the most precise, reliable AI detection results on the market”. Five computer-science conferences and three universities have signed up to use it so far, says the firm’s chief technical officer, Alex Cui, and others are piloting it.

These AI detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, independent analysts say, although no tool can be perfect. And they are “good for screening out places that are pumping out slop”, says Tim Requarth, who studies science communication at New York University’s Langone Health centre in New York City.

But they still sometimes make mistakes, so the tools can be used only as starting points for investigation — and their results are less illuminating for AI-edited writing, in which human and AI text intertwines and there is no clear boundary for problematic use. Nature has tested the tools and interviewed experts to assess how well they do in various situations: where they work, and where they fall short.

How to tell AI writing apart

Until Pangram and others came along, AI-detection software was notoriously unreliable, says Marzena Karpinska, a computer scientist at Simon Fraser University in Burnaby, Canada, who has conducted independent analyses of the tools. Most software analysed statistical characteristics of a text: AI writing tended to display more ‘perplexity’, an estimate of how predictable each word in a sequence is, and less ‘burstiness’, which measures changes in sentence lengths and structures across a text. Some tools would pick up on stylistic quirks in AI writing, such as over-use of the ‘It’s not X, it’s Y’ construction.

But these approaches would too often falsely flag human-authored text as AI-generated. “We didn’t have much faith in the detection,” says Karpinska. Some universities became so fed up with capricious AI detectors that they ended up banning or discouraging staff from using the software.

Then, in early 2024, Spero and Bradley Emi — both computer scientists who had worked at various technology firms since studying together at Stanford University in California — reported a different approach.

They gathered millions of human-written texts and asked LLMs to create ‘mirrors’ of the texts: requesting, for instance, an essay of the same title, length and tone as a human example. Then, they trained machine-learning models to distinguish between the human-written content and the AI mirror, refining the model’s performance by feeding the most challenging cases back in again. Like many such models, the resulting software is a black box. It learns to pick up subtle features of AI writing, often particular to specific LLMs. But these features cannot always be easily described in human terms.

 

How much of the scientific literature is generated by AI?

 

In their 2024 study (which has not been peer reviewed)2, Spero and Emi reported a 0.02% false-positive rate: that is, the model rarely labelled human writing as AI. Karpinska and other independent researchers tested the models3. “We were actually quite surprised that it was working really well. Whatever we threw at it, it was really, really good at detection,” she says.

This July, Epoch AI, a research firm in San Francisco, California, reported that Pangram had zero false positives on 495 human-written texts. Pangram’s own latest technical paper reports a 0.0041% false-positive rate on English texts, with similarly low rates in more than 100 languages4. Tests in the study also show that Pangram doesn’t discriminate against writers who are not fluent English-speakers, a problem flagged with earlier software.

By the time Pangram’s early work became public, GPTZero had also switched from measuring perplexity and burstiness to deploying a similar form of machine learning. It, too, performed well in Karpinska’s study and scored zero false positives on Epoch AI’s test. In February, GPTZero marketed a 0.08% false-positive rate in internal tests, although its technical paper5 more cautiously states this as “sub 1%” across all domains of writing.

AI arms race

The tools’ impressive ability to spot solely human-authored work comes with a trade-off. To avoid flagging human text as AI, both firms accept a higher proportion of false negatives — that is, erroneously clearing some AI-generated text as human. Epoch AI found that Pangram and GPTZero flagged nearly all AI-generated passages made using basic prompts.

 

But when AI was asked to mimic a particular author, some 8% of the resulting passages passed as human.

Karpinska’s study examined a ‘humanizer’ tool — one that rewrites AI-generated text to remove some signs of AI. She found that when detectors were asked to classify human texts versus humanized AI-written text, their false-negative and false-positive rates rose.

But the firms say these tests are already out of date. For instance, Karpinska’s study found that the just-released o1 LLM from OpenAI tripped up an old version of Pangram. Both these tools have now been superseded, and AI-detection firms continually update their models as new versions of LLMs are released.

In the past year, for instance, GPTZero has released 23 updates to its models. Meanwhile, Pangram Labs released its next-generation model, Pangram 4, in July, promising huge advances over its predecessor, including improvements in spotting AI-humanizer signatures. It says its false-negative rate is now only 0.34%, and when AI is asked to imitate styles (as in Epoch’s test), the false-negative rate falls to 2.9%. As with all models, however, performance drops with very short AI passages (of under 50 words).

The upshot is that only the accuracy rates announced by the firms from internal testing can be truly current — but these are not externally verified. “Third-party evaluations will always lag behind the latest detectors,” Requarth says.

The mixed-AI challenge

Both Pangram and GPTZero make money by selling subscriptions for regular or heavy use, but allow some limited free checks. Both have also launched browser tools that check for AI-written text on social media, other web pages and Google docs.

As more writers have experienced being flagged by the software, it’s become clear that the biggest challenge lies in how to approach cases of AI-assisted writing, which is becoming increasingly common. Writers who use AI told Nature that it would be useful to draw a line between using AI to polish or edit human-authored drafts, which they saw as mostly acceptable, and generating a draft with AI from scratch then editing or humanizing it.

Both Pangram and GPTZero break down texts into parts and try to show the degree to which a document is lightly polished with AI, heavily AI-assisted or entirely AI-generated (see ‘How two AI-detectors present their scores’). But in practice, the tools don’t always satisfyingly distinguish between these cases.

 

For instance, in June, Elena Vicario, director of research integrity at the publisher Frontiers, headquartered in Lausanne, Switzerland, wrote a post for The Scholarly Kitchen, a website that posts views on scientific publishing, arguing that AI can be used to support peer review. When Nature ran this article through Pangram in early July, it came up as 100% AI; after Pangram 4 was released, this changed to 96% AI.

But Vicario says her first draft and the ideas that went into it were “entirely human-created” and that she used AI as a tool to polish the text “which is simply best practice”. Editors at The Scholarly Kitchen add that the article was further edited there, so it couldn’t be wholly AI.

If Pangram labels a work as 100% AI, this doesn’t actually mean every word was AI-generated, Spero says. The tool divides a text into segments and judges whether each segment is probably AI-generated, human-written or ‘mixed’. Then, Pangram assigns an overall score on the basis of the proportion of segments flagged. 100% AI means only that each segment was judged as probably AI, even if the segment has some human-written content.

This chunking process means that a paragraph judged ‘human’ in isolation could switch to AI when combined with other text in a segment, Spero adds — explaining why some writers find that sentences extracted from an article can score differently than when judged in the article as a whole.

“If an editor or author used an AI program to smooth out a sentence or clarify the argument during the editing process, we do not view this as a problem,” Vicario says.

Pangram 4, the software launched this July, has changed evaluations in part because it divides text into more fine-grained sections. Whereas the older version analysed segments of some 200–300 words — meaning that a 1,000-word post might end up being divided into only three sections — the new version can analyse chunks as small as 30–40 words.

Spero says that Pangram 4 better distinguishes between a light AI polish — which usually retains a human-authored label — and heavy AI edits. “It takes major AI input like fully-generated sentences or major rewrites to trigger Pangram 4,” he says.

GPTZero analyses texts slightly differently. Like Pangram, it divides them into segments, but at the end, it gives a score that represents its confidence about the whole text, rather than a breakdown of the results of each segment5. Paid users can see which sentences contributed most to GPTZero’s assessment of a work as probably human, AI or mixed. The tool produces 100% confidence that Vicario’s post was AI, described as meaning the software is “highly confident this text is AI-generated”.

 

 

After Nature looked into the case, editors at The Scholarly Kitchen noted on the post that AI had been used “as an editing tool”.

Further questions about Pangram’s evaluations arose when, during the reporting of this article, Spero sent over an examination of articles on Nature’s website. Some articles in a sample from the websites of Nature India and Nature Africa came up as mixed AI–human, and in some cases, 100% AI, he noted, using Pangram 3.3’s assessment.

An examination by the sites’ editorial teams showed that some of the articles flagged were AI-assisted translations of another article or short research highlights that were explicitly written with the aid of AI and had been human-edited. But some were articles written by contributors. They said that they had used AI only to help transcribe or translate interviews, organize notes and edit their drafts. All of these articles went through further rounds of human editing.

In many cases, Pangram 4 judged these articles as merely mixed human–AI, whereas the older version had called them wholly AI. But the newer software still labelled two news reports as 100% AI that contained information derived from original reporting, including interviews. Both sites permit AI use with human oversight, in line with guidance for publications in the Nature Portfolio that says AI should be declared if used for extensive copy-editing or writing, but doesn’t need to be declared for polishing or refining language, although transparency is encouraged. After examining the Pangram results, site editors added notes to some articles to flag the AI assistance.

Overall, as with The Scholarly Kitchen case, the software correctly spotted AI involvement. But Pangram and GPTZero sometimes disagreed on how they rated the articles and on which segments had the strongest indications of being human or AI. And the examples suggest that articles flagged as 100% AI or near 100% can include text written and edited by humans.

Spero says that although Pangram is good at spotting cases in which wholly AI-written paragraphs are inserted between stretches of human-authored text, “there’s a lot of room for improvement” in judging real-world cases in which AI and human writing mingles in a more homogeneous way. For instance, he says, small edits to such documents can sometimes shift Pangram’s score by a large amount — a problem known as ‘jitter’. Pangram’s technical study also notes that when the firm tested using consumer AI to “substantially modify” human-written student essays, the model still labelled the results fully human 41% of the time.

Requarth says that he doesn’t put a lot of faith in exact percentages given by detectors for cases in which AI assistance or editing is involved, and particularly not down to the level of arguing that individual sentences or short chunks of text were AI-written.

The priority for AI detectors should be determining whether a text is fully generated by AI, because “that’s what people are looking for”, says Cui. Determining the extent of AI assistance in a text “is a harder problem”, he says.

A witch hunt?

When Substack readers started to see Pangram’s assessment of posts on the platform, some authors got angry. Sam Illingworth, who studies AI literacy at Edinburgh Napier University, UK, called it a “witch hunt” in a Substack post that Pangram scored as 100% AI. He also ran his text though a humanizer, which switched Pangram’s evaluation to 100% human.

Illingworth says that he makes heavy use of AI tools to edit his drafts. “Absolutely, my work is AI-assisted, but is it 100% AI-generated? No,” he says. The example again shows the potential for confusion with Pangram’s 100%-AI-generated label.

Pangram 4, which was released after Illingworth’s post, rates the original post as 95% AI and 5% human. It assesses the humanized version as 60% AI, “a mix of AI and human-written content” — suggesting that the newer tool does spot traces of AI even in humanized work.

 

 

But posts have appeared online suggesting that some humanizing tools can fool Pangram 4, too. “This cat-and-mouse game on both sides is going to continue,” says Nikhil Garg, a computer scientist at Cornell Tech in New York City.

Illingworth’s concerns remain. He says that letting readers check scores on Substack posts might penalize people whose natural writing style is more AI-like (a concern raised in particular by some neurodivergent writers), and punishes those who use AI to improve work if they don’t have English as a first language — a worry echoed by Evanko, who has been testing Pangram on peer reviews.

“The biggest deficiency is when reviewers used LLMs to translate their reviews from non-Indo-European languages into English,” Evanko says. “The result comes back as fully AI generated.” But the AACR is just bringing in Pangram 4, he adds; in early trials, around one-fifth of reviews classed as 100% AI by the older software have now had their AI fraction reduced, with the most extreme reductions for non-Indo-European reviewers.

Spero says that mere translation shouldn’t trigger an AI flag; he thinks that in the cases Evanko mentions, translation and AI editing are occurring simultaneously.

Cui says that the detector’s score shouldn’t be taken as the result in its own right — but just as a signal for further investigation. A pattern of heavy AI use is more revealing than a judgement on a single article, Spero adds.

The inescapable limitation of AI detection, however good it gets, researchers say, is that although software can spot whether AI was involved in a piece, it can’t prove how it was used or judge what’s ethically acceptable.

Readers really want to know whether a piece was meaningfully authored by a human or whether it was low-effort AI slop or spam, but even a perfect AI-spotter can’t prove that case, says Renée DiResta, a researcher at Georgetown University in Washington DC, who studies scams, disinformation campaigns and other examples of online manipulation and abuse. “We end up surfacing what is easiest to detect rather than addressing the deeper underlying concern,” wrote diResta in a blog post about Pangram’s detector on Substack.

Rising distrust

The new tools are also arriving in an online environment awash with poor-quality AI detection. Inferior services continue to circulate, sometimes offering to circumvent AI detectors by humanizing work, and contributing to mistrust of the software. Many critiques of AI detection reference poor-quality tools.

For instance, Mark Carrigan, who works on digital education at the University of Manchester, UK, and has written a book about generative AI for academics, posted online in July that he’d put his old PhD thesis from 2014, written long before LLMs existed, through a tool called TextGuard, which told him that 62% of the file had signs of AI and offered to humanize it.

He concluded that AI detectors couldn’t be trusted; others have reported similar experiences with services such as TextGuard. But when Nature tested Carrigan’s thesis and other texts, GPTZero and Pangram always correctly identified the older content as human-written. Confusingly, TextGuard’s AI score report includes a graphic stating that it is ‘double-checked’ with GPTZero; Cui says GPTZero has no partnership or association with TextGuard. A spokesperson for TextGuard’s creator, the Hong Kong-based firm Level Media Limited, told Nature that this doesn’t mean the firm doesn’t verify results using GPTZero, but declined to clarify further. In general, the spokesperson wrote, TextGuard bases its assessment on similarity to stylistic patterns and sentence structures that are frequently found in AI writing, and false positives can occur.

 

 

Pangram’s Spero points out that GPTZero itself was acquired in June by the AI-productivity firm Superhuman, formerly known as Grammarly, which offers AI writing assistance, including tools to humanize AI text. Pangram is the leading independent firm that is not yoked to a humanizer operation, he says.

GPTZero’s Cui responds that the tools are separate and not in conflict; GPTZero even detects whether Superhuman’s products were used. For his part, he argues that Pangram is overconfident in its false-positive rates and that it is spuriously precise in the way that it declares that a particular percentage of content is AI or human. Spero disagrees.

As researchers increasingly use AI to write or edit content and shy away from disclosing it, it’s likely that more publishers will adopt AI screening. Already, the manuscript-screening service Proofig AI in Rehovot, Israel, and the research-integrity firm Clearskies in London are offering Pangram to their users. Of research publishers that replied to Nature’s queries, Science journals say they recently incorporated iThenticate’s AI text-detection feature, and MDPI says that it has built an in-house AI detector named Binoculars. Springer Nature says it is exploring in-house and third-party AI-detection tools. (Nature’s news team is editorially independent of its publisher, Springer Nature.)

In another development, this month, US firm Anthropic said it would introduce watermarking into its Claude AI models to comply with the European Union’s AI Act — although it’s unclear how easily editing could remove the marks, or whether writers will avoid Claude as a result.

Ultimately, it’s likely that if watermarking and AI-detection tools become widespread, writers, including scientists, will need to start more transparently disclosing how they use AI, and to find ways to prove or document their process, says Garg. And that might help everyone find answers to the underlying issue: what counts as ethically acceptable AI use?” [1]

1. AI-detection tools have made huge leaps forward — how good are they? Nature 656, 808-811 (2026) By Miryam Naddaf & Richard Van Noorden