ArticlesAI and ProductivityOver a third of web pages written by AI since ChatGPT launched, Pew finds

Over a third of web pages written by AI since ChatGPT launched, Pew finds

Bar chart showing the proportion of web pages written by AI across .com, .org, .edu and .gov domains since ChatGPT launched

The human web is shrinking faster than most people realised

The majority of what you read online may no longer have been written by a person. A new Pew Research Center report, published on 20 August 2026, found that more than one-third of web pages written by AI have been pushed live since the launch of ChatGPT, showing clear signs of machine authorship, and the trend is accelerating fastest in exactly the corners of the web where UK consumers spend most of their time.

What the Pew report actually found

Pew researchers pulled nearly half a million English-language web pages from the Common Crawl web archive, covering roughly the past five years. They ran the text through an AI detection tool called Open Pangram, which weighs statistical patterns across large batches of text rather than flagging individual words or phrases.

Across the full sample, around 10% of pages showed significant signs of AI authorship. That figure sounds modest until you account for the age of the pages. Many in the dataset predate AI writing tools entirely, which naturally pulls the average down. When researchers filtered to content published after ChatGPT’s release, the proportion showing signs of AI authorship rose to 35%.

It is worth being precise about what “signs of AI authorship” means here. The researchers are not claiming that one-third of every sentence was typed by a machine. Detection models can misclassify pages in both directions, and plenty of the flagged content was likely AI-assisted rather than fully generated. Even with that caveat, the scale of the shift is significant.

The domain breakdown tells the real story

Not all corners of the web are changing at the same rate, and the pattern of where AI content is concentrated matters a great deal for how you use the web day-to-day.

Commercial .com domains showed signs of AI authorship at roughly ten times the rate found on .edu or .gov domains, both of which registered around 1%. Nonprofit .org domains fell in between at around 4.6%. Before ChatGPT launched, AI-authorship indicators appeared at broadly similar rates across all the main top-level domains. The divergence is entirely a post-ChatGPT phenomenon.

For UK readers, that gap is directly relevant. The .com space is where most commercial content lives: product reviews, price comparison write-ups, affiliate guides, news commentary, opinion pieces, and the kind of “best of” articles that show up when you search for anything from a new boiler to a business bank account. The domains you trust most for neutral information are far less affected. The ones you rely on for practical purchasing research are where AI content is most concentrated.

Bots writing for bots

The Pew findings land alongside a separate milestone that compounds the concern. Internet infrastructure provider Cloudflare recently confirmed that bot web traffic has now overtaken human web traffic, a threshold the company says was reached sooner than its own estimates had predicted.

Put those two data points together and the picture shifts. More of the content on the web is being produced by AI, and more of the traffic reading that content is automated too. The human web, in the sense of pages written by people and read by people, is now a minority activity by volume.

The Pew study period runs to July 2026, so this is not a retrospective look at early adoption. It reflects what is being published and read right now.

AI in journalism too

Open Pangram, the same detection tool used in the Pew study, has also appeared in separate research finding AI-generated text in roughly 9% of US newspaper articles this year, including opinion pages at major outlets. That figure is worth pausing on. Opinion journalism has historically been one of the formats most dependent on individual voice and judgement. Its inclusion in these findings suggests the shift is not confined to low-quality content farms.

For UK readers who rely on online journalism to form views on politics, business, or consumer decisions, this is a reasonable prompt to think more carefully about provenance. The byline and the outlet name are not the guarantees of human authorship they once were.

Will AI detection get better?

Reliably identifying AI text is currently imperfect. Detection models work on statistical patterns, and those patterns can produce false positives and false negatives at the individual page level. It is an inherently probabilistic exercise, which is partly why Pew was careful to frame its findings in terms of “signs of” AI authorship rather than definitive attribution.

The medium-term picture may change, though. Anthropic and other large AI developers are already working on model-level text fingerprinting. If adopted widely, this approach would make AI-generated text easier to identify at source, rather than relying on reverse-engineered pattern detection after the fact. That would significantly reduce the error rate in both directions. Whether it would also slow the growth of AI content is a separate question entirely.

What this means for you

The practical effect of these findings depends on how you use the web. If you primarily use it for transactional research, comparing products, reading reviews, or evaluating suppliers, the concentration of AI content on .com domains means you are already operating in the highest-density zone. That does not mean everything you read is unreliable, but it does mean the bar for cross-referencing claims is now higher than it was two or three years ago.

For small business owners and tradespeople, this has a specific edge to it. A significant share of the content ranking for commercial search queries, including guides on software, tools, and services relevant to running a business, is now at least partially AI-generated. The information may be broadly accurate, or it may be plausible-sounding and wrong. Checking a second or third source, especially one on a less commercially driven domain, is increasingly worth the extra two minutes.

There is also a search engine dimension worth watching. Google and other search engines have been adjusting how they evaluate content quality for several years, and the volume of AI-generated pages now hitting the web is a direct challenge to those systems. How rankings evolve in response will affect which content surfaces first, particularly for commercially competitive queries. If search engines begin deprioritising AI-heavy content, the sites that invested in human writing may find themselves better positioned. If they do not, the incentive to publish AI content at volume only grows stronger.

The web is still useful, just differently

The rise of AI-generated content does not make the web useless, but it does change the skills needed to use it well. Source literacy, the habit of asking who wrote something and why, was always worth having. It is now closer to essential.

Frequently asked questions

Does the Pew study prove that AI-written content is lower quality or inaccurate?

Not directly. The study identifies signs of AI authorship but does not assess accuracy or quality. AI-generated content can range from accurate and useful to plausible-sounding but wrong. The concern is that volume-driven AI publishing reduces the editorial checks that would normally catch errors, not that AI authorship is inherently a quality signal on its own.

Should I trust .gov and .edu websites more now?

Based on the Pew findings, those domains show AI-authorship indicators at around 1%, which is dramatically lower than .com domains. That does not mean every government or university page is perfect, but the rate of AI-assisted content is far lower, and those domains typically have stronger editorial oversight.

Will search engines start labelling AI-generated content?

There is no widespread labelling system in place yet. Google has indicated it evaluates content on quality and usefulness rather than method of production. Some AI developers, including Anthropic, are working on text fingerprinting that could make AI content easier to identify at source, but this has not yet been deployed at scale.

How does Open Pangram detect AI-written text?

Open Pangram does not flag individual words or phrases. It analyses statistical patterns across large volumes of text, looking for distributional signals that differ from typical human writing. Detection at the individual page level can produce errors in both directions, so the Pew findings are more reliable as aggregate trends than as verdicts on any single page.

Does this affect UK websites specifically, or is it a US-focused finding?

The Pew study used English-language pages from the Common Crawl archive, which is a global dataset. The .com domain in particular is used internationally, so the findings apply broadly to the English-language web including UK commercial sites. There is no reason to think UK .com domains are significantly different from the global trend.

The web has always had low-quality content, but the speed and scale at which AI is now producing it represents something different in kind, and the Pew findings give that shift a number for the first time.