For years, the “dead internet theory” lived on the fringes: the half-joking claim that most of what we read online is no longer written by people. It was always more vibe than measurement. Now researchers have attached a number to a version of it.
According to a new study, 31.1% of the text in a quality-filtered sample of the web from August 2026 was labeled as AI-generated, up from 27.5% just two months earlier. That’s not the whole internet, and the number comes from an AI detector rather than a confession from every publisher. But it’s one of the most careful attempts yet to measure how much of the web’s useful-looking text now comes from machines.
The more surprising finding is what that means for AI itself. You might expect more AI-written text to make training the next generation of models easier, since there’s simply more data around. The researchers found the opposite. Past a certain point, AI-generated text in the training mix makes models worse at understanding human writing, and compensating for it costs significantly more computing power.
This article explains what the study found, why AI text makes training harder, how AI labs are responding, and what it all means for search, publishers, and anyone who creates content online.
What the Research Found
The study, titled “How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text,” was posted to arXiv on September 30, 2026. It comes from researchers at Pangram Labs, a company that makes AI-text detection software, and the University of Maryland. AI Weekly was among the first outlets to cover it.
Measuring the AI share
The team took web crawl data and applied FineWeb quality filtering, a widely used method for cleaning up raw web text before it’s used to train language models. Then they ran Pangram’s AI detector over the filtered text. Their results:
- June 2026: 27.5% of tokens were labeled as AI-generated.
- August 2026: 31.1% of tokens were labeled as AI-generated.
A token is a small chunk of text, roughly a word or part of a word. So this measures the share of text, not the share of web pages or websites.
The jump of 3.6 percentage points in two months is notable. If that pace continued, AI-labeled text would make up a much larger share of quality-filtered web data within a year or two, though trends like this rarely move in a straight line.
Testing the effect on training
The second, more ambitious part of the study tested what this AI text does to model training. The researchers trained 800 small and mid-sized language models, ranging from about 20 million to nearly 1 billion parameters, while varying the ratio of AI-generated to human-written text in their training data. They then measured how well each model predicted held-out human text.
They also released their data and tools publicly, including WildAI, an 83-billion-token corpus labeled for AI origin, topic, and format, plus all 800 trained models and their code. That openness lets other researchers check the work.
Important caveats
Before sharing the 31% figure, keep a few things in mind:
- It’s not the whole web. The figure applies to a quality-filtered crawl, the kind of text AI labs typically train on. It doesn’t cover video, images, social media posts, or private platforms.
- It depends on a detector. AI-text detectors make mistakes, and no detector is perfect. The study’s authors work for a detection company, which has a commercial interest in this topic, though their data and code are public.
- It’s a preprint. The paper hasn’t yet gone through formal peer review.
- Other studies use different methods. An Ahrefs analysis of 900,000 new web pages in April 2025 found that 74.2% contained some AI-generated content, but only 2.5% were classified as purely AI. The numbers differ because Ahrefs counted pages containing any AI text, while this study measured the share of all tokens.
The broad conclusion is consistent across studies: a large and growing share of what’s published online now involves AI writing.
Why AI Text Makes Training More Expensive
The headline number from the training experiments is 1.6x. At the August 2026 level of 31.1% AI text, the researchers estimate a model needs about 1.6 times the computing power to reach the same quality on human text as a model trained only on human writing. That estimate applies to a common training setup of about 20 tokens of data per model parameter.
To put that in perspective, compute is the largest cost in building frontier AI models. A 60% increase in compute for the same result is a serious tax, especially at a time when training runs can cost hundreds of millions of dollars or more.
A little helps, then it hurts
The relationship isn’t simple, and that’s what makes the study interesting. According to the paper’s abstract:
- For models short on data, adding some AI text at first improves performance on human text. More data, even imperfect data, helps a hungry model.
- But the benefit levels off quickly, and as more AI text is added, it reverses and starts to do harm.
- For models already trained on plenty of human text, adding AI text makes things worse almost immediately, while the same amount of fresh human text would keep improving the model.
In short, AI-written text is a poor substitute for human writing when the goal is to understand how humans write.
Why would AI text hurt?
The paper focuses on measuring the effect rather than fully explaining it, but researchers have long suspected a few reasons. AI-generated text tends to be more uniform than human writing. It favors common phrasings, predictable structures, and safe word choices. Human writing is messier, more varied, and full of unusual expressions, local knowledge, and specific detail.
A model trained heavily on AI text learns the patterns of AI writing rather than the full range of human language. It gets better at predicting machine prose and worse at predicting the long tail of how real people write.
Not quite model collapse, but related
This connects to an earlier concern called model collapse. In a 2024 study published in Nature, researchers showed that models trained repeatedly on their own outputs gradually lose information about rare events and unusual data, until quality degrades badly.
The new study makes a different point. It looks at “wild” AI text: content produced by many different models, written for human readers, and mixed unlabeled into ordinary web data. That’s the real-world situation AI labs actually face. The finding suggests you don’t need an extreme feedback loop to see harm. Ordinary web pollution is enough to impose a measurable cost.
Standard scaling laws miss it
AI labs plan their training runs using scaling laws, formulas that predict how a model’s quality improves with more data and compute. The best-known are the “Chinchilla” scaling laws from DeepMind researchers in 2022.
The new study found that Chinchilla-style laws fail to predict how AI text affects training. The authors propose a revised formula with separate terms for the benefit and harm of AI text, allowing the value of an AI token to switch from positive to negative. When fitted on smaller models, their formula predicted results for models up to 3.6 times larger with 41% lower error than the best existing law.
The practical warning: labs that plan training runs using standard formulas on today’s web data may underestimate how much compute they need and overestimate how good their models will be.
What AI Labs Are Doing About It
If AI text in training data carries a cost, the obvious response is to find more human writing and less machine writing. That’s easier said than done, but several strategies are already in play.
1. Filtering out AI text
The study’s first recommendation is to filter AI-generated text when the goal is a model that understands human writing. Detection tools can flag likely AI content so it can be removed or down-weighted. The trade-off is that detectors make mistakes, and filtering too aggressively can throw away good human writing along with the machine-made material.
2. Reusing human text before adding AI text
The researchers also suggest that labs repeat their existing human-written data, training on it more than once, before expanding their datasets with AI-heavy web text. Repetition has its own limits, but the findings suggest familiar human data can be worth more than fresh AI-written data.
3. Valuing older data
Text published before generative AI went mainstream in late 2022 is far less likely to contain AI writing. Some researchers compare it to “low-background steel,” the metal produced before nuclear testing that’s prized for sensitive instruments because it isn’t contaminated by radiation. Older crawls and archives may become more valuable as a clean reference.
4. Licensing human content directly
Over the past few years, AI companies have signed licensing deals with news publishers, online communities, and other content owners to access high-quality human writing. Findings like this strengthen the economic case for those deals. If human text is genuinely scarcer and more valuable, the people and organizations who produce it gain bargaining power.
5. Deliberate synthetic data, not accidental
There’s an important distinction between “wild” AI text scraped from the web and synthetic data that labs generate on purpose. Labs increasingly create carefully designed synthetic data for specific tasks like math, coding, and reasoning, where answers can be checked. The study itself notes that AI text remains useful when the target is AI-style text. The problem is unlabeled, unvetted machine text mixed in with everything else.
6. Measuring human and AI performance separately
The researchers recommend that labs report how their models perform on human text and AI text separately. A model can look good on a blended test while quietly getting worse at understanding real human writing. Separate measurements make that trade-off visible.
7. Provenance and labeling
Longer term, the industry is working on ways to label AI content at the source. Approaches include invisible watermarks embedded in AI-generated text and images, and content credential standards that record how a piece of content was created. These tools are still far from universal, and text watermarks can be weakened by editing, but wider adoption would make filtering far more reliable.
What It Means for the Web: Spam, Search, and Trust
The training-cost finding is a problem for AI labs. But a web where nearly a third of quality text is machine-written affects everyone who uses it.
Content farms got cheaper
Generative AI made it nearly free to produce plausible articles at huge scale. That’s been a gift to content farms and SEO spammers, who publish thousands of pages targeting search keywords with little original value. Search engines have responded. Google, for example, introduced a spam policy in 2024 against “scaled content abuse,” targeting pages produced in bulk mainly to manipulate rankings, whether by humans, AI, or both.
But it’s a moving target. The cost of producing passable content keeps falling, and the line between AI-assisted and AI-generated writing keeps blurring.
A feedback loop in search
AI is now on both sides of search. AI writes many of the pages, and AI summarizes them in search results. Ahrefs found that 86.5% of top-ranking pages contain some AI-generated content, and its analysis suggested Google’s AI Overviews may lean toward citing AI-assisted pages. When AI summaries draw on AI-written sources, errors and generic framing can be recycled and amplified without a human ever checking the original facts.
Sameness is the subtle cost
Even when AI-written content is accurate, it tends to sound alike. If a growing share of articles on a topic are produced by similar models with similar prompts, the web loses variety: fewer unusual viewpoints, fewer first-hand experiences, less local detail. That’s the same quality that makes AI text less useful for training. It also makes the web less useful for readers looking for something new.
Trust becomes the scarce resource
As machine text grows, readers will rely more on signals of who stands behind content: named authors with real expertise, publications with a reputation to protect, first-hand evidence, and transparent sourcing. Platforms that can show where content came from may earn more trust. Anonymous, generic pages may earn less, from readers and from search engines alike.
What It Means for Content Creators
If you publish content regularly, this research raises an uncomfortable question. When nearly a third of quality web text is AI-written, does original human work become more valuable, or does it just get buried faster? The honest answer is both, and which one wins depends on what you publish.
The case that human work gains value
The study’s core finding is an economic signal: genuinely human writing is now worth more than machine text to the companies building AI. That’s showing up in licensing deals between AI firms and publishers, and it may extend further as labs hunt for clean data.
Readers value it too. In a sea of interchangeable articles, content with a distinct voice, original reporting, first-hand testing, or real expertise stands out. AI can summarize what’s already been published; it can’t interview your sources, test a product in your hands, or share your actual experience.
The case that it gets buried
Volume still matters. A single creator can’t out-publish a content farm generating thousands of pages a day. If search engines and social feeds can’t reliably tell original work from polished imitation, good content can lose visibility to sheer quantity. And AI search summaries may answer readers’ questions without them ever clicking through to the original source.
What to do about it
Whatever tools you use in your workflow, a few practices help your work stand out in a machine-heavy web:
- Add something only you can add. Original data, first-hand testing, interviews, case studies, local knowledge, or a clear personal point of view. If an AI could produce your article from existing pages, so can everyone else.
- Use AI as an assistant, not an author. AI can help with research, outlines, and editing. The value comes from the human judgment, facts, and perspective layered on top.
- Check every fact. AI-written pages recycle each other’s errors. Verifying claims against primary sources is a real differentiator.
- Show who you are. Clear author bylines, credentials, and sourcing signal trust to readers and search engines.
- Build direct audiences. Newsletters, communities, and loyal followers are less exposed to search algorithm swings and AI summaries.
- Be thoughtful about how your work is used. Review your site’s settings for AI crawlers and any licensing opportunities, depending on whether you want your content used for training.
The creators most at risk are those producing generic, interchangeable content, because that’s exactly what AI does cheaply. The creators best placed are those whose work carries something a model can’t fake.
Final Thoughts
The 31% figure is a striking data point, but the more important idea is the one behind it: AI-generated text isn’t a free substitute for human writing. Past a point, it makes AI models worse at understanding people, and fixing that costs real money. The “just scale up” approach to building better models now faces a data-quality constraint that standard formulas don’t account for.
That has knock-on effects well beyond AI labs. Human-written content is becoming a scarcer and more valuable resource. Search engines face a harder job separating originality from volume. And creators face a choice between competing with machines on quantity, which they’ll lose, or on authenticity and insight, where they still have the edge.
The study has limits. It’s a preprint, it relies on an AI detector, and it measures one slice of the web. But the trend it captures is consistent with other research, and it’s moving fast. The web’s next few years will be shaped by how well humans, platforms, and AI companies learn to tell the difference between what people wrote and what machines generated.
Frequently Asked Questions
Is 31% of the internet really AI-generated?
Not exactly. The study found that 31.1% of tokens in a quality-filtered web crawl from August 2026 were labeled AI-generated by Pangram’s detector. That’s a specific slice of web text used for AI training, not the entire internet, and detectors can make mistakes.
Why does AI-generated text make training more expensive?
AI text tends to be more uniform than human writing. When too much of it is mixed into training data, models get worse at predicting human text. The researchers estimate that at today’s levels, matching the quality of human-only training needs about 1.6 times the compute in a common training setup.
Is this the same as model collapse?
It’s related but different. Model collapse describes models degrading when trained repeatedly on their own outputs. This study looked at ordinary AI text from many models mixed into web data, and found measurable harm even without that extreme feedback loop.
How are AI companies responding?
Approaches include filtering AI-generated text out of training data, reusing human-written data, valuing older pre-2022 data, licensing content from publishers, generating carefully designed synthetic data for specific tasks, and developing watermarking and content labeling standards.
Does this make human-written content more valuable?
In many ways, yes. Original, expert, first-hand content is scarcer and more useful, both to readers and to AI developers. But visibility isn’t guaranteed, because mass-produced content can still crowd search results and feeds.












