• Home
  • Contact Us
  • Free SEO Tools
Newsletter
PostDune
  • Business
    • Economics
    • Finance
    • Marketing
  • Entertainment
  • Fashion
  • Health
  • Home Improvement
  • Politics
  • Sports
  • Technology
  • Travel
No Result
View All Result
  • Business
    • Economics
    • Finance
    • Marketing
  • Entertainment
  • Fashion
  • Health
  • Home Improvement
  • Politics
  • Sports
  • Technology
  • Travel
No Result
View All Result
PostDune
No Result
View All Result
Home Technology

Nearly a Third of Quality Web Text Is Now AI-Written, and It’s Making AI Harder to Train

Daisy by Daisy
October 9, 2026
in Technology
0
AI-Generated Text Challenges Training

AI-Generated Text Challenges Training

189
SHARES
1.5k
VIEWS
Share on FacebookShare on Twitter

For years, the “dead internet theory” lived on the fringes: the half-joking claim that most of what we read online is no longer written by people. It was always more vibe than measurement. Now researchers have attached a number to a version of it.

According to a new study, 31.1% of the text in a quality-filtered sample of the web from August 2026 was labeled as AI-generated, up from 27.5% just two months earlier. That’s not the whole internet, and the number comes from an AI detector rather than a confession from every publisher. But it’s one of the most careful attempts yet to measure how much of the web’s useful-looking text now comes from machines.

The more surprising finding is what that means for AI itself. You might expect more AI-written text to make training the next generation of models easier, since there’s simply more data around. The researchers found the opposite. Past a certain point, AI-generated text in the training mix makes models worse at understanding human writing, and compensating for it costs significantly more computing power.

This article explains what the study found, why AI text makes training harder, how AI labs are responding, and what it all means for search, publishers, and anyone who creates content online.

What the Research Found

The study, titled “How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text,” was posted to arXiv on September 30, 2026. It comes from researchers at Pangram Labs, a company that makes AI-text detection software, and the University of Maryland. AI Weekly was among the first outlets to cover it.

Measuring the AI share

The team took web crawl data and applied FineWeb quality filtering, a widely used method for cleaning up raw web text before it’s used to train language models. Then they ran Pangram’s AI detector over the filtered text. Their results:

  • June 2026: 27.5% of tokens were labeled as AI-generated.
  • August 2026: 31.1% of tokens were labeled as AI-generated.

A token is a small chunk of text, roughly a word or part of a word. So this measures the share of text, not the share of web pages or websites.

The jump of 3.6 percentage points in two months is notable. If that pace continued, AI-labeled text would make up a much larger share of quality-filtered web data within a year or two, though trends like this rarely move in a straight line.

Testing the effect on training

The second, more ambitious part of the study tested what this AI text does to model training. The researchers trained 800 small and mid-sized language models, ranging from about 20 million to nearly 1 billion parameters, while varying the ratio of AI-generated to human-written text in their training data. They then measured how well each model predicted held-out human text.

They also released their data and tools publicly, including WildAI, an 83-billion-token corpus labeled for AI origin, topic, and format, plus all 800 trained models and their code. That openness lets other researchers check the work.

Important caveats

Before sharing the 31% figure, keep a few things in mind:

  • It’s not the whole web. The figure applies to a quality-filtered crawl, the kind of text AI labs typically train on. It doesn’t cover video, images, social media posts, or private platforms.
  • It depends on a detector. AI-text detectors make mistakes, and no detector is perfect. The study’s authors work for a detection company, which has a commercial interest in this topic, though their data and code are public.
  • It’s a preprint. The paper hasn’t yet gone through formal peer review.
  • Other studies use different methods. An Ahrefs analysis of 900,000 new web pages in April 2025 found that 74.2% contained some AI-generated content, but only 2.5% were classified as purely AI. The numbers differ because Ahrefs counted pages containing any AI text, while this study measured the share of all tokens.

The broad conclusion is consistent across studies: a large and growing share of what’s published online now involves AI writing.

Why AI Text Makes Training More Expensive

The headline number from the training experiments is 1.6x. At the August 2026 level of 31.1% AI text, the researchers estimate a model needs about 1.6 times the computing power to reach the same quality on human text as a model trained only on human writing. That estimate applies to a common training setup of about 20 tokens of data per model parameter.

To put that in perspective, compute is the largest cost in building frontier AI models. A 60% increase in compute for the same result is a serious tax, especially at a time when training runs can cost hundreds of millions of dollars or more.

A little helps, then it hurts

The relationship isn’t simple, and that’s what makes the study interesting. According to the paper’s abstract:

  • For models short on data, adding some AI text at first improves performance on human text. More data, even imperfect data, helps a hungry model.
  • But the benefit levels off quickly, and as more AI text is added, it reverses and starts to do harm.
  • For models already trained on plenty of human text, adding AI text makes things worse almost immediately, while the same amount of fresh human text would keep improving the model.

In short, AI-written text is a poor substitute for human writing when the goal is to understand how humans write.

Why would AI text hurt?

The paper focuses on measuring the effect rather than fully explaining it, but researchers have long suspected a few reasons. AI-generated text tends to be more uniform than human writing. It favors common phrasings, predictable structures, and safe word choices. Human writing is messier, more varied, and full of unusual expressions, local knowledge, and specific detail.

A model trained heavily on AI text learns the patterns of AI writing rather than the full range of human language. It gets better at predicting machine prose and worse at predicting the long tail of how real people write.

Not quite model collapse, but related

This connects to an earlier concern called model collapse. In a 2024 study published in Nature, researchers showed that models trained repeatedly on their own outputs gradually lose information about rare events and unusual data, until quality degrades badly.

The new study makes a different point. It looks at “wild” AI text: content produced by many different models, written for human readers, and mixed unlabeled into ordinary web data. That’s the real-world situation AI labs actually face. The finding suggests you don’t need an extreme feedback loop to see harm. Ordinary web pollution is enough to impose a measurable cost.

Standard scaling laws miss it

AI labs plan their training runs using scaling laws, formulas that predict how a model’s quality improves with more data and compute. The best-known are the “Chinchilla” scaling laws from DeepMind researchers in 2022.

The new study found that Chinchilla-style laws fail to predict how AI text affects training. The authors propose a revised formula with separate terms for the benefit and harm of AI text, allowing the value of an AI token to switch from positive to negative. When fitted on smaller models, their formula predicted results for models up to 3.6 times larger with 41% lower error than the best existing law.

The practical warning: labs that plan training runs using standard formulas on today’s web data may underestimate how much compute they need and overestimate how good their models will be.

What AI Labs Are Doing About It

If AI text in training data carries a cost, the obvious response is to find more human writing and less machine writing. That’s easier said than done, but several strategies are already in play.

1. Filtering out AI text

The study’s first recommendation is to filter AI-generated text when the goal is a model that understands human writing. Detection tools can flag likely AI content so it can be removed or down-weighted. The trade-off is that detectors make mistakes, and filtering too aggressively can throw away good human writing along with the machine-made material.

2. Reusing human text before adding AI text

The researchers also suggest that labs repeat their existing human-written data, training on it more than once, before expanding their datasets with AI-heavy web text. Repetition has its own limits, but the findings suggest familiar human data can be worth more than fresh AI-written data.

3. Valuing older data

Text published before generative AI went mainstream in late 2022 is far less likely to contain AI writing. Some researchers compare it to “low-background steel,” the metal produced before nuclear testing that’s prized for sensitive instruments because it isn’t contaminated by radiation. Older crawls and archives may become more valuable as a clean reference.

4. Licensing human content directly

Over the past few years, AI companies have signed licensing deals with news publishers, online communities, and other content owners to access high-quality human writing. Findings like this strengthen the economic case for those deals. If human text is genuinely scarcer and more valuable, the people and organizations who produce it gain bargaining power.

5. Deliberate synthetic data, not accidental

There’s an important distinction between “wild” AI text scraped from the web and synthetic data that labs generate on purpose. Labs increasingly create carefully designed synthetic data for specific tasks like math, coding, and reasoning, where answers can be checked. The study itself notes that AI text remains useful when the target is AI-style text. The problem is unlabeled, unvetted machine text mixed in with everything else.

6. Measuring human and AI performance separately

The researchers recommend that labs report how their models perform on human text and AI text separately. A model can look good on a blended test while quietly getting worse at understanding real human writing. Separate measurements make that trade-off visible.

7. Provenance and labeling

Longer term, the industry is working on ways to label AI content at the source. Approaches include invisible watermarks embedded in AI-generated text and images, and content credential standards that record how a piece of content was created. These tools are still far from universal, and text watermarks can be weakened by editing, but wider adoption would make filtering far more reliable.

What It Means for the Web: Spam, Search, and Trust

The training-cost finding is a problem for AI labs. But a web where nearly a third of quality text is machine-written affects everyone who uses it.

Content farms got cheaper

Generative AI made it nearly free to produce plausible articles at huge scale. That’s been a gift to content farms and SEO spammers, who publish thousands of pages targeting search keywords with little original value. Search engines have responded. Google, for example, introduced a spam policy in 2024 against “scaled content abuse,” targeting pages produced in bulk mainly to manipulate rankings, whether by humans, AI, or both.

But it’s a moving target. The cost of producing passable content keeps falling, and the line between AI-assisted and AI-generated writing keeps blurring.

A feedback loop in search

AI is now on both sides of search. AI writes many of the pages, and AI summarizes them in search results. Ahrefs found that 86.5% of top-ranking pages contain some AI-generated content, and its analysis suggested Google’s AI Overviews may lean toward citing AI-assisted pages. When AI summaries draw on AI-written sources, errors and generic framing can be recycled and amplified without a human ever checking the original facts.

Sameness is the subtle cost

Even when AI-written content is accurate, it tends to sound alike. If a growing share of articles on a topic are produced by similar models with similar prompts, the web loses variety: fewer unusual viewpoints, fewer first-hand experiences, less local detail. That’s the same quality that makes AI text less useful for training. It also makes the web less useful for readers looking for something new.

Trust becomes the scarce resource

As machine text grows, readers will rely more on signals of who stands behind content: named authors with real expertise, publications with a reputation to protect, first-hand evidence, and transparent sourcing. Platforms that can show where content came from may earn more trust. Anonymous, generic pages may earn less, from readers and from search engines alike.

What It Means for Content Creators

If you publish content regularly, this research raises an uncomfortable question. When nearly a third of quality web text is AI-written, does original human work become more valuable, or does it just get buried faster? The honest answer is both, and which one wins depends on what you publish.

The case that human work gains value

The study’s core finding is an economic signal: genuinely human writing is now worth more than machine text to the companies building AI. That’s showing up in licensing deals between AI firms and publishers, and it may extend further as labs hunt for clean data.

Readers value it too. In a sea of interchangeable articles, content with a distinct voice, original reporting, first-hand testing, or real expertise stands out. AI can summarize what’s already been published; it can’t interview your sources, test a product in your hands, or share your actual experience.

The case that it gets buried

Volume still matters. A single creator can’t out-publish a content farm generating thousands of pages a day. If search engines and social feeds can’t reliably tell original work from polished imitation, good content can lose visibility to sheer quantity. And AI search summaries may answer readers’ questions without them ever clicking through to the original source.

What to do about it

Whatever tools you use in your workflow, a few practices help your work stand out in a machine-heavy web:

  • Add something only you can add. Original data, first-hand testing, interviews, case studies, local knowledge, or a clear personal point of view. If an AI could produce your article from existing pages, so can everyone else.
  • Use AI as an assistant, not an author. AI can help with research, outlines, and editing. The value comes from the human judgment, facts, and perspective layered on top.
  • Check every fact. AI-written pages recycle each other’s errors. Verifying claims against primary sources is a real differentiator.
  • Show who you are. Clear author bylines, credentials, and sourcing signal trust to readers and search engines.
  • Build direct audiences. Newsletters, communities, and loyal followers are less exposed to search algorithm swings and AI summaries.
  • Be thoughtful about how your work is used. Review your site’s settings for AI crawlers and any licensing opportunities, depending on whether you want your content used for training.

The creators most at risk are those producing generic, interchangeable content, because that’s exactly what AI does cheaply. The creators best placed are those whose work carries something a model can’t fake.

Final Thoughts

The 31% figure is a striking data point, but the more important idea is the one behind it: AI-generated text isn’t a free substitute for human writing. Past a point, it makes AI models worse at understanding people, and fixing that costs real money. The “just scale up” approach to building better models now faces a data-quality constraint that standard formulas don’t account for.

That has knock-on effects well beyond AI labs. Human-written content is becoming a scarcer and more valuable resource. Search engines face a harder job separating originality from volume. And creators face a choice between competing with machines on quantity, which they’ll lose, or on authenticity and insight, where they still have the edge.

The study has limits. It’s a preprint, it relies on an AI detector, and it measures one slice of the web. But the trend it captures is consistent with other research, and it’s moving fast. The web’s next few years will be shaped by how well humans, platforms, and AI companies learn to tell the difference between what people wrote and what machines generated.

Frequently Asked Questions

Is 31% of the internet really AI-generated?

Not exactly. The study found that 31.1% of tokens in a quality-filtered web crawl from August 2026 were labeled AI-generated by Pangram’s detector. That’s a specific slice of web text used for AI training, not the entire internet, and detectors can make mistakes.

Why does AI-generated text make training more expensive?

AI text tends to be more uniform than human writing. When too much of it is mixed into training data, models get worse at predicting human text. The researchers estimate that at today’s levels, matching the quality of human-only training needs about 1.6 times the compute in a common training setup.

Is this the same as model collapse?

It’s related but different. Model collapse describes models degrading when trained repeatedly on their own outputs. This study looked at ordinary AI text from many models mixed into web data, and found measurable harm even without that extreme feedback loop.

How are AI companies responding?

Approaches include filtering AI-generated text out of training data, reusing human-written data, valuing older pre-2022 data, licensing content from publishers, generating carefully designed synthetic data for specific tasks, and developing watermarking and content labeling standards.

Does this make human-written content more valuable?

In many ways, yes. Original, expert, first-hand content is scarcer and more useful, both to readers and to AI developers. But visibility isn’t guaranteed, because mass-produced content can still crowd search results and feeds.

Share76Tweet47
Previous Post

Anthropic’s Reported IPO Plans: A Trillion-Dollar Valuation on Top of Record Losses

  • Trending
  • Comments
  • Latest
How to Stop a Metal Bed Frame from Squeaking

How to Stop a Metal Bed Frame from Squeaking: Ultimate Guide

February 1, 2025
Most Popular Sports in Europe

The Most Popular Sports in Europe

January 21, 2025
5 Things To Consider When Visiting A Display Home In Kellyville

5 Things To Consider When Visiting A Display Home In Kellyville

December 19, 2025
Largest Economies in the World in 2026 - Top 10 Countries by GDP Ranking

Largest Economies in the World in 2026: Top 10 Countries by GDP

August 31, 2026
How AI Is Changing User Search Behavior in 2026

How AI Is Changing User Search Behavior in 2026

9
Google vs. AI Search in 2026

Why Users Are Switching from Google Search to AI Chatbots

5
Rise of Zero-Click Searches

The Rise of Zero-Click Searches: What It Means for Websites

2
Traditional search vs AI assistant comparison

The Complete Guide to AI Search in 2026: How Artificial Intelligence Is Transforming the Future of Search

2
AI-Generated Text Challenges Training

Nearly a Third of Quality Web Text Is Now AI-Written, and It’s Making AI Harder to Train

October 9, 2026
Anthropic's Reported IPO Plans

Anthropic’s Reported IPO Plans: A Trillion-Dollar Valuation on Top of Record Losses

October 5, 2026
Best Back-End Development Companies for Enterprise Software in 2026

Best Back-End Development Companies for Enterprise Software in 2026

October 5, 2026
Dune Networks

Dune Networks: The Israeli Chip Startup Behind Broadcom’s StrataDNX

October 3, 2026
PostDune

Categories

  • AI
  • Automotive
  • Beauty
  • Business
  • Digital Marketing
  • Economics
  • Education
  • Entertainment
  • Fashion
  • Finance
  • Gaming
  • General
  • Health
  • Home Improvement
  • Lifestyle
  • Marketing
  • News
  • Real Estate
  • Sports
  • Technology
  • Travel

Recent Posts

  • Nearly a Third of Quality Web Text Is Now AI-Written, and It’s Making AI Harder to Train
  • Anthropic’s Reported IPO Plans: A Trillion-Dollar Valuation on Top of Record Losses
  • Best Back-End Development Companies for Enterprise Software in 2026
  • Dune Networks: The Israeli Chip Startup Behind Broadcom’s StrataDNX
  • Top AI-Assisted Software Development Companies to Watch in 2026
  • Home
  • Privacy Policy
  • Disclaimer
  • Write for us
  • Terms and conditions
  • Contact Us
  • Free SEO Tools

Copyright © 2026 by postdune.com. All Rights Reserved.

No Result
View All Result
  • Home
  • Business
    • Economics
    • Finance
    • Marketing
  • Entertainment
  • Fashion
  • Health
  • Home Improvement
  • Politics
  • Sports
  • Technology
  • Travel
  • Contact Us

Copyright © 2021 by postdune.com. All Rights Reserved.

This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.