More than a third of web pages published since ChatGPT’s release carry signs of artificial intelligence authorship

Tech and AI

More than a third of web pages published since ChatGPT's release carry signs of artificial intelligence authorship

By Staff Writer  |  21 August 2026

Close-up of black printed text on a textured paper page, with one word in sharp focus

An analysis of 490,000 English-language web pages taken from a public web archive found that 10 per cent of all pages in a July 2026 sample show strong signs of AI authorship, and that the share rises above one third once pages written before ChatGPT was released are removed. The rate is highest on .com addresses and around ten times lower on .edu and .gov.

The method is worth stating before the numbers, because it decides what the numbers mean. Researchers drew almost half a million English-language pages from the past five years out of the Common Crawl web archive, starting a couple of years before ChatGPT was released in November 2022, and ran the text through a detection model that scores writing for patterns more common in machine output than in human output. The claim being made is that a page shows signs of AI authorship, which is not the same claim as a page having been written entirely by a machine.

Ten per cent of everything, a third of everything new

In a random sample of 10,000 pages collected in July 2026, 10 per cent carried strong signs of AI authorship. That figure is held down by the age of the web itself, since a large share of any random sample was written before the tools existed. Filter to pages published after ChatGPT was released and the share passes one third.

Where it lands is uneven, and that is the finding a professional reader should take away. In the 2026 samples about one page in ten on a .com address carries the signs, against 4.6 per cent on .org and roughly one per cent on both .edu and .gov. When ChatGPT first appeared the four top-level domains sat at similar rates. They have separated since.

The tells, and why any single one proves nothing

The analysis also counts the surface features that have become more common as machine-written text has spread. Compared with a 2023 snapshot of the web, the em dash now appears about twice as often. Oxford commas are up 63 per cent. A set of 27 words the models reach for more readily than people do has more than doubled in use. The construction that sets up a comparison by first denying one half of it has nearly tripled, although it remains rare in absolute terms.

AI models tend to use these dashes a lot more than humans typically do. They're also more likely than human authors to list items in threes and to use Oxford commas in lists.

Samuel Bestvater, senior data scientist, Pew Research Center

The explanation offered is that models learn the patterns of the writing they were trained on, that some kinds of writing were over-represented in that training, and that the habits of those texts come back out the other end. Journalistic and academic prose is heavy on the em dash, so the machines are heavy on it too.

None of those features settles anything about a single document. Human writers use all of them. The analysis says so directly, and adds that detection models sometimes misclassify individual documents in both directions. What the model is doing is reading subtler statistical patterns in word choice and sentence structure across a very large body of text, and it is at that scale, not at the level of one page, that the result stands up.

What it changes for anyone who commissions writing

Two practical points follow. The first is that a house style which bans a handful of punctuation marks and a list of words is now, whether or not it was meant to be, a partial machine-text filter, and firms that have adopted one have accidentally acquired a screening tool. The second is the harder one. If the commercial web is drifting towards a shared register at this rate, the value of a page that plainly did not come off a production line goes up, and the cost of proving that it did not goes up with it.

For anyone relying on the open web as a source of fact, the number to hold on to is the last one. Official and academic addresses are still running at about one per cent. That gap is the reason a primary record is worth the extra hour it takes to find.