A new estimate suggests that linguistic traces associated with large language models spread through English-language biomedical publishing at extraordinary speed. In 2023, the estimated share was 19 percent, close to one paper in five. Across 2025 it was 77 percent. For December 2025 alone, it reached 89 percent, or nearly nine papers in ten.
Those figures come from an August 2026 preprint by Lena Holzwarth, Rita González-Márquez and Dmitry Kobak. They are not a count of authors who admitted using AI, nor the output of a detector accusing individual papers. The study estimates a corpus-wide vocabulary signal. It cannot tell whether a tool translated a sentence, polished grammar or generated pages of prose.
That boundary matters. The result is still remarkable, but “linguistic traces” and “AI-written papers” are not interchangeable claims.
What the 89 percent actually covers
The team analyzed 1,194,287 English-language papers dated from 2017 through 2025 in the PubMed Central Open Access Subset. Each paper needed recognizable sections such as an introduction, methods, results and discussion, with enough text in every section to support the analysis.
This design improves on studying titles or abstracts alone because it can compare writing across a full paper. It does not cover every biomedical publication. It covers open-access full texts deposited in one repository and meeting the structural filters. About 1.8 percent of the included records were preprints, and the authors note that a preprint and its published version can both be present.
The headline numbers therefore describe this large, selective corpus. Generalizing them to all biomedical papers, every language or every publisher would go beyond the evidence.
The method models a missing counterfactual
No archive contains an alternate 2025 in which ChatGPT was never released. The researchers had to estimate what scientific vocabulary would have looked like without LLM assistance.
They started with 379 nontechnical style words identified in earlier work as unusually common after ChatGPT appeared. Examples include ordinary terms such as “these,” “potential” and “delves.” For each word, they fitted a pre-ChatGPT trend using papers from 2018 through 2022, projected that trend forward, then measured how far actual usage rose above the projection.
The statistical estimator asks what minimum share of documents would need LLM assistance to produce the excess. In simulations created under its assumptions, its error stayed below two percentage points. The real corpus, however, has no verified label showing which authors used which tools. Simulation checks the mathematics of the model; it is not independent ground-truth validation.
The central assumption is that pre-2023 human vocabulary trends would otherwise have continued. A separate shift in scientific style could inflate the estimate. Humans copying conspicuous AI phrasing could also count as part of the signal, while authors deliberately avoiding those words could suppress it.
Longer papers are easier to flag
The chance of seeing at least one marker increases as a document gets longer. That is why the authors report both whole-section estimates and comparisons based on equal 255-word samples.
When all sections of each paper were combined, the estimate rose from 19 percent in 2023 to 52 percent in 2024, 77 percent across 2025 and 89 percent in December 2025. For a random 255-word passage from a full paper, the December estimate was 40 percent.
Equal-length samples also exposed differences between sections. In December 2025, the estimated signal was 68 percent in discussions, 67 percent in abstracts, 59 percent in introductions, 46 percent in results and 32 percent in methods.
That pattern is consistent with authors using language tools more often for summaries and interpretation than for procedural details. It does not prove that explanation, because vocabulary alone cannot reconstruct a writer’s workflow. The gap between 89 percent and 40 percent is the more secure lesson: document length materially changes the headline estimate.
A second study finds the rise, but not the same level
An independent peer-reviewed PNAS study by Kyle Siler examined 7.3 million full-text journal articles published by Elsevier, Frontiers, MDPI and PLOS from 2020 through 2025. Using 228 focal words and a different modeling strategy, it estimated LLM diffusion at 12 percent in 2023 and 57 percent in 2025.
Those values are below the new preprint’s 19 percent for 2023 and 77 percent for all of 2025. That is not necessarily a contradiction. The studies use different publishers, inclusion rules, word lists and estimators. Their convergence is about direction and speed, not one settled percentage.
Together they make a stronger case that the rise is real. The distance between their estimates is also a warning against treating any single number as a direct measurement of undisclosed AI use.
The COVID comparison is about vocabulary, not harm
The claim that this vocabulary shift exceeded the disruption caused by COVID-19 comes from the team’s earlier peer-reviewed Science Advances study. That analysis tracked 15.1 million English-language PubMed abstracts from 2010 through 2024.
It found 454 words used more often than expected in 2024. At the 2021 peak of the pandemic vocabulary shock, it found 190. After related inflections were reduced to unique lemmas, the comparison was 343 versus 180.
The two shifts had different characters. COVID altered what scientists were studying, so its excess vocabulary was dominated by content nouns such as “pandemic” and “coronavirus.” The post-ChatGPT shift was dominated by style: of 379 excess style words, 66 percent were verbs and 14 percent were adjectives.
“Bigger than COVID” therefore means broader by this word-count measure. It does not compare social damage, research quality or public-health consequences. Nor does it imply that papers containing these words are fraudulent.
A trace does not tell us what the AI did
The 89 percent analysis is a preprint and has not completed peer review. It also cannot identify individual assisted papers. That makes it unsuitable for screening authors, investigating misconduct or attaching an AI label to a manuscript.
LLM assistance covers very different practices. A researcher writing in a second language might use a tool for translation or grammar. Another might ask it to reorganize an argument. A third might generate unsupported claims or references. The vocabulary model places those possibilities inside one population-level signal and cannot separate them.
I think the most honest reading is that AI-shaped language has probably become normal in biomedical publishing, while its exact prevalence remains model-dependent. The next useful evidence will need verified disclosure records, publisher workflow data or provenance systems that connect a particular text change to a particular tool.
Vocabulary can show that the culture of scientific writing changed unusually fast. It cannot, by itself, tell us whether that change improved access, concealed responsibility or affected the reliability of any one paper.