AI Writes 20% of Scientific Papers: How Not to Fall Into the "AI-Generated" Trap in 2026
The statistics, the detection tools, the consequences — and a co-author checklist for staying on the right side of the line.
The number arrived in the research literature this year with the quiet force of something that had been obvious in retrospect. For computer science, review preprints containing AI-generated text increased from about 7% in 2023 to 43% in 2025. Non-review manuscripts in this field that contained AI-generated text also grew from around 3% to 23% during the same period. In biomedical research, the trajectory is different in scale but identical in direction: the proportion of published articles including AI-generated text increased from 0% in January 2022 to 11.3% in March 2025, with a significant increasing trend throughout.
A 2025 survey by the journal Nature found that 57% of scientists admitted to seeking writing help from AI in the last two years, and 72% said they wanted to do so in the next two years.
The headline that has been circulating — that AI now writes approximately 20% of scientific papers — is, depending on the field and the definition of "AI-generated," either an underestimate or an overestimate. What is not in dispute is the direction: the proportion of scientific text with detectable AI involvement is rising rapidly across every field measured, and the infrastructure being built to detect and respond to that involvement is keeping pace with a speed that many researchers have not yet registered. The tools are real, the accuracy is improving, and the consequences for undisclosed AI use are increasingly concrete.
This article is about navigating that landscape — not by avoiding AI, which would be both impractical and unnecessary, but by understanding what the current detection environment looks like, what journals are doing with it, and how to use AI assistance in manuscript preparation in ways that are transparent, compliant, and safe for a research career.
What the Data Actually Shows
The statistics on AI in scientific writing require careful interpretation because the measurement tools, the definitions, and the fields vary significantly. Researchers applied an LLM-detection model to the abstracts and introductions of 1,121,912 preprints and journal-published papers from January 2020 to September 2024 on arXiv, bioRxiv, and 15 Nature portfolio journals. The analysis revealed a sharp uptick in LLM-modified content just months after the release of ChatGPT in November 2022 — appearing so quickly that it means people were using AI for writing immediately. The biggest increases were in areas closest to AI research itself.
A study scanning nearly 7,000 manuscript abstracts submitted to the journal Organization Science between January 2021 and February 2026, along with some 8,000 peer-review reports, found that submission volume has risen by 42% since November 2022. The increase is partially attributable to AI assistance reducing the friction of writing — making it faster and cheaper to produce manuscripts, and therefore producing more of them.
Approximately 20% of ICLR conference reviews and 12% of Nature Communications reviews were classified as AI-generated in 2025. A Nature survey of 1,600 academics found that more than 50% have used AI tools while peer reviewing manuscripts, even when the core judgments remained human-authored.
The peer review data is in some ways more alarming than the manuscript data — not because peer reviewers using AI for assistance is inherently problematic, but because Nature Portfolio policies explicitly prohibit uploading manuscripts to generative AI systems during review, and ask reviewers who used AI in any way to declare that use transparently. A peer reviewer who uploads an unpublished manuscript to an AI system is violating confidentiality requirements regardless of whether they disclose the AI assistance, because the manuscript belongs to the submitting authors and has not been consented to third-party processing.
The Detection Infrastructure Journals Are Actually Using
The gap between what researchers assume about AI detection and what journals are deploying is wide, and it is closing rapidly. The assumption that AI detection tools are unreliable — based largely on early experiences with tools like GPTZero and Turnitin that had significant false positive rates, particularly for non-native English writers — is increasingly outdated.
Pangram, the detection tool being adopted by major publishers including the American Association for Cancer Research, operates on overlapping windows of text rather than whole documents, allowing it to flag localised AI-written segments embedded within otherwise human-authored manuscripts. In head-to-head benchmarks it reaches approximately 99–99.8% accuracy with false-positive rates on the order of 0.01–0.1% across domains, including in scientific writing and essays by non-native English speakers, and remains highly effective even against adversarial "humaniser" paraphrases explicitly optimised to evade detection.
Independent evaluation found Pangram to be the only commercially available AI detection tool to satisfy a strict false positive rate cap of 0.005 without sacrificing accuracy, with near-zero false negative rates robust across models, threshold rules, ultra-short passages, and humaniser tools. The 2025 evaluation found its false positive rate to be less than 0.001 when evaluating samples written by GPT-4.1, Claude Opus 4, Claude Sonnet 4, and Gemini 2.0 Flash.
This matters for a specific and underappreciated reason: the argument that AI detection tools are too unreliable to be used fairly against authors is losing its empirical basis. A tool with a false positive rate of 0.001 applied to a manuscript with ten sections does not meaningfully risk wrongly accusing a human writer. The standard author response — "detection tools are unreliable, I didn't use AI, this is just my writing style" — is becoming harder to sustain as the tools' accuracy improves and their adoption by major publishers accelerates.
Elsevier, JAMA Network, Science, PLOS ONE, and Nature now all require authors to disclose any AI use, ban AI as an author, and/or permit AI only with explicit disclosure requirements. Strong disclosure requirements remain one of the most effective safeguards publishers can implement.
What Undisclosed AI Use Actually Looks Like
The cases of AI-related retraction that have accumulated since 2023 fall into three distinct categories, each representing a different kind of risk.
-
Wholesale generation
Manuscripts substantially or entirely written by an AI system, with no meaningful human intellectual contribution, submitted without disclosure. These are the cases most likely to be caught by detection tools and most likely to result in immediate retraction when discovered. Two separate retraction notices in February 2025 addressed AI-assisted papers on rare gliomas — while the data may have been sound, the language and references had been manipulated using AI without transparent authorship or vetting, undermining trust even in the scientific content.
-
Hallucinated citations
References to papers that do not exist, generated by AI systems that produce plausible-looking but fabricated bibliographic entries. These are detectable through reference verification tools that major journals increasingly deploy as part of automated pre-submission screening. A manuscript with even one hallucinated reference that passes peer review and is subsequently identified is a retraction risk regardless of whether the rest of the paper is entirely accurate.
-
Partial AI generation without disclosure
Manuscripts where a researcher used AI assistance for substantial portions of the writing — introduction, discussion, conclusion — while conducting the underlying science themselves. This is where the majority of the current AI-in-science volume sits, and where the rules are clearest but the practice most widespread. Using AI to write a significant portion of a manuscript without disclosing it is, under the policies of every major journal, a form of misrepresentation — even if the science behind it is entirely the researcher's own.
The False Positive Problem for Non-Native English Writers
There is a genuine equity concern that has been somewhat obscured by the improving accuracy of detection tools: the residual false positive rate, while low, falls disproportionately on researchers writing in English as a second or third language.
Earlier detection tools were documented to flag non-native English writing as AI-generated at significantly higher rates than native English writing — because the structural and stylistic patterns that AI systems produce overlap with some of the patterns produced by people writing carefully and formally in a language that is not their own. The best current tools, including Pangram, have been specifically evaluated on non-native English scientific writing and report false positive rates low enough that this bias is substantially reduced. But it is not eliminated.
For a researcher in Kyiv, Tbilisi, or Almaty whose scientific English is excellent but structured differently from native production, this residual risk is real. The most effective mitigation is not to avoid AI assistance — it is to disclose AI use accurately and specifically, which removes the basis for a false positive accusation entirely. A researcher who discloses that they used an AI tool to improve the clarity of a section they wrote in full cannot be accused of undisclosed AI generation, regardless of what any detection tool reports.
The Co-Author Checklist for 2026
The following checklist is designed for research teams preparing manuscripts for submission to peer-reviewed journals in 2026. It is not a checklist for avoiding AI — it is a checklist for using AI correctly, disclosing it accurately, and protecting every author on the paper from consequences they did not knowingly incur.
What every research group needs before submission
Inventory — what AI was used and where
- Every use of AI identified — writing assistance, literature summary, translation, figure generation, data analysis.
- Every AI tool identified by name and version (ChatGPT-4o, Claude Sonnet 4, Grammarly).
- Every section where AI generated or substantially revised text identified specifically.
- Every reference verified against its actual source — not the citation as generated by an AI tool.
Disclosure — what the journal requires
- The target journal's specific AI disclosure policy reviewed and located.
- AI use in writing assistance disclosed in the Acknowledgements; AI use in data analysis or figures disclosed in Methods.
- The disclosure specifies which tools were used and for what specific purpose.
- The corresponding author has confirmed no hidden AI use — including in sections contributed by co-authors.
Verification — what needs human confirmation
- Every factual claim verified against primary sources rather than AI-generated summaries.
- Every reference checked to confirm it exists, is accessible, and supports the specific claim attributed to it.
- The manuscript read in its entirety by at least one author who did not use AI in their sections.
- Any AI-generated passages substantially revised — not merely accepted as produced.
Co-author responsibility
- Every named co-author aware that AI was used in manuscript preparation.
- Every co-author has reviewed and taken responsibility for their sections, including any AI-assisted ones.
- No co-author has inadvertently used an AI tool that uploaded the manuscript to a third-party system — including during peer review.
- Any author using AI for language assistance is prepared to confirm this in the disclosure rather than leave it undisclosed.
Detection self-check
- The manuscript run through a credible AI detection tool before submission — not to game detection, but to identify sections where AI assistance is more detectable than realised.
- Any flagged sections accurately described in the disclosure.
- For any false positive on human-written text, documentation of the writing process (drafts, notes, version history) retained to support a response if challenged.
The Direction of Travel
The trajectory of AI in scientific writing is one-directional and accelerating. AI now accounts for 5.8%–8.8% of scientific research output depending on the field, up from below 1% in 2010. The detection infrastructure is improving in accuracy, expanding in deployment, and becoming integrated into editorial workflows at major publishers as a routine step rather than an exceptional investigation.
The researchers who will navigate this environment most successfully are not those who avoid AI — the productivity advantages are too significant for avoidance to be realistic at the field level — and not those who use it covertly, hoping detection will not keep pace. They are the researchers who use AI transparently, disclose it accurately, verify everything it produces, and treat the disclosure not as a confession but as a demonstration of the intellectual honesty that distinguishes their work from the paper-mill outputs that made the detection infrastructure necessary in the first place.
The trap is not using AI. The trap is using it as if no one is watching — in a year when, demonstrably, everyone is.
Questions readers ask after this piece
How much of scientific writing now involves AI-generated text?
It varies sharply by field. In computer science, review preprints containing AI-generated text rose from about 7% in 2023 to 43% in 2025. In biomedical research, the proportion of published articles including AI-generated text rose from 0% in January 2022 to 11.3% in March 2025. A 2025 Nature survey found 57% of scientists had sought writing help from AI in the previous two years. The headline figure of around 20% is, depending on field and definition, either an underestimate or an overestimate — but the direction is not in dispute.
Are AI detection tools accurate enough to be used against authors?
Increasingly, yes. Pangram, adopted by major publishers, reaches approximately 99–99.8% accuracy with false-positive rates of 0.01–0.1% across domains, and remains effective against adversarial humaniser paraphrases. Independent evaluation found a false positive rate below 0.001 against GPT-4.1, Claude Opus 4, Claude Sonnet 4, and Gemini 2.0 Flash. The old argument that detection tools are too unreliable to be used fairly is losing its empirical basis.
Which journals now require AI disclosure?
Elsevier, JAMA Network, Science, PLOS ONE, and Nature all now require authors to disclose any AI use, ban AI as an author, and/or permit AI only with explicit disclosure. Strong disclosure requirements are considered one of the most effective safeguards publishers can implement, and they are increasingly enforced through automated pre-submission screening.
What kinds of undisclosed AI use lead to retraction?
Three categories. Wholesale generation — manuscripts substantially written by AI with no meaningful human contribution and no disclosure — is most likely to be caught and retracted. Hallucinated citations — references to papers that do not exist — are detectable through reference verification tools. Partially AI-generated text without disclosure, where a researcher used AI for substantial portions while doing the science themselves, is the most widespread category and is treated as misrepresentation under every major journal's policy even when the science is sound.
Do AI detection tools unfairly flag non-native English writers?
Earlier tools did, at significantly higher rates, because careful formal writing in a second language overlaps stylistically with AI output. The best current tools, including Pangram, have been evaluated specifically on non-native English scientific writing and report false positive rates low enough that this bias is substantially reduced — though not eliminated. The most effective mitigation is to disclose AI language assistance accurately and specifically, which removes the basis for a false-positive accusation entirely.
Can a peer reviewer use AI to help review a manuscript?
With significant caution. Nature Portfolio policies explicitly prohibit uploading manuscripts to generative AI systems during review and require reviewers to declare any AI use transparently. A reviewer who uploads an unpublished manuscript to an AI system violates confidentiality regardless of whether they disclose the AI assistance, because the manuscript belongs to the submitting authors and has not been consented to third-party processing. Roughly 20% of ICLR and 12% of Nature Communications reviews were classified as AI-generated in 2025.