Skip to content
The Big Picture August 29, 2026 Updated August 29, 2026

35% of Web Pages Written Since ChatGPT Show AI Authorship

Pew classified roughly 500,000 pages with the detector Substack uses. Commercial .com domains run 10 times the .edu rate, which is exactly where your content lives.

By The State of AI Marketing newsroom
Share
Editorial illustration for: 35% of Web Pages Written Since ChatGPT Show AI Authorship
Credit: JAC Growth Marketing

Your marketing site sits in the most AI-written part of the internet, and now there’s a number on it. Pages on .com domains show signs of AI authorship at roughly ten times the rate of .edu and .gov pages, which means “we publish original content” has quietly become a claim your buyer has reason to doubt and you have no way to prove.

That comes from Pew Research Center, published August 20. Pew classified nearly 500,000 English-language pages pulled from the Common Crawl archive between January 2021 and July 2026, then took a random sample of 10,000 pages from July 2026 and ran them through Open Pangram, the same detector Substack adopted to catch machine-written newsletters.

The headline numbers: 10% of all pages sampled in July 2026 show significant signs of AI authorship, rising to 35% among pages carrying a post-ChatGPT publication date. Split by domain, that’s 10% on .com against 4.6% on .org and about 1% on .edu and .gov.

Nearly every number published about AI content until now came from a company selling AI detection or AI writing. Pew sells neither. That’s the rarest thing in this story.

The tells are the ones you have been told not to trust

Pew also tracked the surface markers, and they moved in the direction everyone suspected. Em dashes went from 5.79 uses per 10,000 words in early 2023 to 11.19 in early 2026. Oxford commas rose 63%. “Delve”, “interplay” and “testament” more than doubled.

Two things to hold at once here. Individually those markers prove nothing, and Pew says so: plenty of careful human writers use em dashes and always have. But the aggregate shift is real and measurable, which is why the detector works on a corpus and not on your competitor’s blog post.

We’ve written our own prose without em dashes since launch for exactly this reason, and the honest version is that the rule is about reader perception rather than truth. It’s the same instinct behind watermarking AI-edited copy, where the mark works as a signal a reader can act on rather than as proof of anything. A dash doesn’t make writing machine-made. It makes a reader suspect it is, and after 2026 that suspicion has data behind it.

The caveat that changes what you can say about it

The most important limitation isn’t in the headline coverage. Matt G. Southern at Search Engine Journal caught it: Pew’s threshold captures AI-assisted editing and AI-generation identically. A paragraph a person wrote and then ran through a grammar tool scores the same as a paragraph a model produced from a prompt.

Nobody’s separated those two, in this study or any other. Google Docs and Microsoft Word both ship native AI editing now, and neither leaves a different fingerprint from ChatGPT.

So the honest reading of 35% is “a third of new pages have been through a model at some point”, not “a third of new pages were written by a robot”. Those are very different claims about the web, and only the weaker one is supported. Anyone citing this number as proof of a content-slop wave is overreading it, and that includes the coverage this week.

Pew’s own framing is the careful one: AI detection models “sometimes misclassify individual documents”, and the finding describes patterns across a large corpus rather than a verdict on any page.

What it does to differentiation

We argued in June that AI made content nearly free to produce and differentiation nearly impossible. That was an argument from mechanism. This is the measurement, and it lands harder than the argument did.

The competitive problem isn’t that AI writing is bad. Much of it is fine. The problem is that fine is now the floor across a third of everything new, and a floor that high erases the middle of the market. Content that’s merely competent no longer signals anything about the company that published it, because competent is what a prompt produces.

What survives is what a model can’t generate from a prompt: your own numbers, your customers by name, a position you’d defend in a room, and work you did that nobody else could have done. Everything else is now indistinguishable from the 35%.

The counter-case: this may say more about the long tail than about you

The strongest argument against reading this as a marketing story is a sampling one.

Common Crawl is dominated by the enormous, low-value tail of the web: parked domains, thin affiliate pages, programmatic listings, scraped aggregators. Those are exactly the pages you’d expect to be machine-written, they vastly outnumber real publisher and brand pages, and .com is where they all live. So the .com-versus-.edu gap could be measuring spam density rather than anything about corporate marketing sites.

Pew doesn’t break the .com figure down by page type, so this can’t be settled from the published data. It’s a real objection and it should temper the number.

It doesn’t rescue you, though. Your buyer doesn’t sample the web. They see search results and answer engines, which draw from that same .com pool. Whether the 10% is spam or strategy, the pool your content competes in reads as machine-written at ten times the rate of the pool nobody is trying to sell in.

The move is to stop treating “we write our own content” as a differentiator and start treating it as an unverifiable claim, because that’s what it now is. Publish the things a prompt can’t produce and put them where a person can check: named authors with real histories, first-party numbers with methods attached, customer specifics you had to earn. Anything you’d be comfortable seeing a competitor generate in an afternoon is no longer worth the afternoon.

Quoted in this story

  • Matt G. Southern, Senior News Writer, Search Engine Journal (source)

Want your perspective in coverage like this? Get quoted.

Sources

This story is part of our running coverage: the full picture →

Get Net Effect.

The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.

More from The Big Picture