AI-Generated Content Now Makes Up 35% of the Internet. That’s Not Even the Scary Part.

Written by

in

Table of contents

AI-generated content already makes up roughly 35% of all new pages published online and that number comes from mid-2025. You just read something. It felt fine. No red flags. No obvious errors. Smooth, coherent, forgettable.

That’s the problem.

The Numbers No One Wants to Sit With

Researchers from Imperial College London, the Internet Archive, and Stanford analyzed millions of web pages. Their finding: roughly 35% of new content published online by mid-2025 was AI-generated content — text produced or heavily assisted by large language models.

Not 5%. Not the spam corners of the web.

One in three new pages across the open internet.

For context: in 2022, estimates put that figure at a fraction of a percent. Three years. That’s how long it took to fundamentally change the composition of the internet invisibly, and without anyone voting on it.

AI-Generated Content

When AI-generated content hits one-third market penetration, it stops being a trend. It becomes infrastructure.

What You Think Is Happening versus What Is Actually Happening

Ask anyone journalists, academics, your LinkedIn feed, what AI-generated content does to the web. You’ll get a consistent story: more misinformation, lower factual accuracy, epistemic bubbles, style homogenization. Everything sounds the same, everything is worse.

Except the data doesn’t back that up.

The researchers tested six distinct hypotheses about the negative effects of AI-generated content on the internet:

  1. Factual accuracy is declining
  2. Writing style is homogenizing
  3. Epistemic isolation (filter bubbles) is increasing
  4. Reader engagement is dropping
  5. Semantic diversity is narrowing
  6. Overall tone is becoming more positive

Four out of six: not confirmed.

This should make you stop. Not because it means AI-generated content is harmless. Because our collective intuition about what’s happening is wrong in exactly the places that matter. And when your threat model is wrong, you’re not protected you’re just confident.

What Specifically Didn’t Hold Up

Facts didn’t get worse. Style didn’t become uniform. Filter bubbles aren’t growing faster than before.

That doesn’t mean these problems don’t exist. It means AI-generated content isn’t their direct, measurable cause at least not at this stage of the research.

The Two Effects That Actually Showed Up

Two hypotheses held under scrutiny.

The first: decreased semantic diversity. This isn’t about every article sounding the same. It’s about the range of concepts, perspectives, and frameworks circulating online getting narrower. AI systems optimize toward a center the statistical average of everything they’ve trained on. So AI-generated content publishes toward that center. Over and over.

The internet is starting to think with one brain.

The second confirmed effect: increased positivity. AI-generated content skews more optimistic than human writing. That sounds harmless. It isn’t. A more positive internet is also a less critical one less willing to call out failure, less capable of saying “this is broken.” The feedback loop that makes information systems self-correct is getting quieter.

Combine those two. You get a medium that speaks with one voice and has a relentlessly upbeat take on everything.

That’s not journalism. That’s not knowledge. That’s a very large content farm with good grammar.

A Concrete Example

You search for a product review. Twenty results all positive, all similarly worded, all missing specific downsides. Not because the product is good. Because AI-generated content defaults away from negative assessments: “positive” articles get better engagement metrics.

The result: the internet stops being a place where you can verify whether something actually works.

How They Actually Measured This

Detecting AI-generated content at scale is genuinely hard. A single detector fails on well-crafted text the better the model, the harder to catch.

The researchers took a different approach:

  • Stacked multiple detection methods simultaneously
  • Cross-validated results across methods
  • Stratified the sample by site type, language, and time period

It’s more rigorous than most “AI content crisis” pieces you’ve read in the news. Which makes the 35% figure harder to dismiss as measurement noise.

A Dark Forecast and Three Things You Can Do

If 35% is the mid-2025 baseline for AI-generated content, where is that curve now?

The internet’s value as a thinking tool came from friction. Different people, different assumptions, different conclusions colliding in public. That friction is how errors got corrected, and blind spots got exposed. Semantic homogenization is friction reduction optimized away not by a conspiracy, but by a thousand content teams all producing AI-generated content, optimizing for the same metrics, converging on the same center.

Three concrete steps:

  1. Diversify sources actively — not through an algorithm. Seek out voices that actively disagree
    with the mainstream.
  2. Search for criticism, not reviews — ask “what’s wrong with X,” not “what is X.”
  3. Go to the primary data — read the report, not the article about the report.

The question isn’t whether AI-generated content is changing the internet anymore.

The question is whether the internet as a medium for collective thought survives what’s already happened to it. And whether you’ll be one of the people who knows how to navigate what’s left.


Want More? Subscribe to The Dossier

Every week in your inbox:

📡 THE INTELLIGENCE FEED – 3-5 curated links: [Research] [Policy] [Tools] [Incidents]
💡 ONE ADVICE – One actionable AI/cybersecurity tip you can use today

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *