AI Model Collapse Isn’t a Supply Problem — It’s Curation

AI model collapse isn't about running out of text — it's a curation failure, and the difference decides which labs keep improving.

6 min read

Every large language model on the market was built by devouring nearly everything humans have ever written down — and that well is running dry. That crunch has pushed AI labs toward a fix, synthetic data, that carries a hidden cost: a failure mode researchers call AI model collapse, where models trained on their own generated output slowly forget the rare, unusual, and minority cases in the world. Epoch AI, the research group that tracks AI compute and data trends, estimates the effective stock of quality-adjusted, publicly available human text usable for training sits at roughly 300 trillion tokens (a token is roughly a word-and-a-half, the basic unit these models are trained and billed on). At current trends, that stock could be exhausted somewhere between 2026 and 2032 — and pulled forward to as early as 2027 if labs keep overtraining, deliberately training past the point where more compute stops paying off.

That countdown-clock framing is the story most coverage tells, and it buries the sharper point. Running out of human text and model collapse aren’t two separate crises — they’re cause and effect. Labs turn to synthetic data because the well is drying up, and it’s the uncurated, recursive use of that synthetic data that erodes a model’s grip on rare cases and diverse viewpoints, quietly, before it ever shows up in a benchmark score. The fix was never going to be “more data” — it’s who controls the mix a model sees, and whether a human is still filtering it.

The AI Training Data Exhaustion Clock Everyone Cites

The exhaustion numbers come from a widely cited 2024 Epoch AI paper, Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data (Villalobos et al.). Their 300 trillion token ceiling is deliberately conservative — it already accounts for quality filtering, deduplication, and the fact that most text on the open internet isn’t useful training material. Two forces are squeezing that ceiling from both sides at once. Training-dataset sizes have grown roughly 2.5x a year, while the compute labs throw at training has grown roughly 4x a year — compute is outrunning its own fuel supply.

Overtraining is the wrinkle that makes the timeline scarier than the ceiling alone suggests. Compute-optimal training balances model size against data volume for the best performance per unit of compute; overtraining means running well past that ratio anyway, because a smaller, cheaper-to-serve model that’s been overtrained can beat a compute-optimal model at the same running cost. Labs already do this routinely — it’s standard practice, not an edge case. Push that dial by a modest 5x factor, Epoch estimates, and the human-text stock could be fully spent by 2027, not 2032.

When Does the Human Text Run Out? Projected exhaustion of the ~300 trillion token effective human-text stock 2026 2027 2028 2029 2030 2031 2032 Baseline window: full exhaustion 2026–2032 With 5x overtraining: as early as 2027 ~300 trillion tokens of usable human text

Source: Epoch AI (Villalobos et al., arXiv:2211.04325)

What AI Model Collapse Actually Does Inside a Model

The mechanism paper behind the term is Ilia Shumailov and colleagues’ AI Models Collapse When Trained on Recursively Generated Data, published in Nature in 2024. The core finding: train a model on its own outputs, or another model’s outputs, across successive generations without an ongoing anchor in real human-generated data, and performance degrades in a specific, predictable direction. Each generation’s outputs sample from an already-narrowed version of the true distribution — the full range of sentences, facts, and styles a language could actually produce — so statistically rare events get sampled less often than they truly occur. The result, replicated by the authors across large language models, variational autoencoders, and Gaussian mixture models alike, is that the tails vanish first.

In practice, distribution tails are the unusual phrasing, the minority dialect, the rare medical presentation, the niche historical fact — anything that shows up in the real world less often than a model’s internal average expects. Early-stage collapse looks subtle: a slight flattening of diversity, a model that answers competently but blandly, nuance quietly filed off before it registers on a benchmark. Late-stage collapse is not subtle at all — outputs converge toward a single repetitive mode, disconnected from the reality the model was meant to represent. Nature issued an Author Correction to the paper in 2025, and a separate analysis from Borji has questioned how far the original numeric results generalize, so the mechanism looks solid while its exact magnitude remains an open argument.

Compute Is Outrunning Its Own Fuel Cumulative relative growth vs. baseline year (log scale) 1x 10x 100x 1000x Yr 0 Yr 1 Yr 2 Yr 3 Yr 4 Yr 5 Yr 6 Relative growth (log scale) Training compute (~4x/year) Dataset size (~2.5x/year)

Source: Epoch AI growth-rate estimates (Villalobos et al., arXiv:2211.04325)

Get the next one by email

Physics, engineering and the people behind them. No spam, unsubscribe any time.

Distillation vs. Model Collapse: Same Synthetic Data, Opposite Outcome

Here’s the detail that undercuts the panic: synthetic data doesn’t reliably break models — uncurated synthetic data does. Distillation is the technique of training a smaller model on a larger model’s carefully filtered outputs, letting the compact model inherit the bigger one’s reasoning without inheriting its running cost. Microsoft’s Phi line of small models, built on the paper Textbooks Are All You Need, made its name doing exactly this — synthetic, textbook-quality examples generated and filtered with real human oversight, not scraped in bulk and recycled blind. It works well enough that distillation is now a standard production tool, not a research curiosity.

The difference between a distilled model that gets sharper and a recursively trained model that collapses isn’t the origin of the text — both are synthetic. It’s whether a human, or a well-designed filter, keeps a live connection to the real distribution: rejecting bland, averaged-out outputs, deliberately preserving edge cases, and mixing in fresh human-generated data instead of looping a model’s output back into itself unchecked. Model collapse is what happens when curation stops, not an inevitable tax on using AI to generate training data. That reframes the data shortage from an unsolvable ceiling into a governance problem labs can actually manage — for now.

⚡ PHOTON’S TAKE

I’ve spent enough time around large-scale data infrastructure to recognize this pattern instantly: recycle output back into input without a filter, and you don’t get more of what you had — you get less of everything except the average. That’s model collapse in one sentence. The industry isn’t actually running out of data; it’s running out of patience for curation, which is cheap to skip now and expensive to fix later. Distillation already proves the fix works. The labs that win the next five years won’t be the ones with the most synthetic tokens — they’ll be the ones who never stopped filtering them.

Who Controls the Mix Decides Whether AI Model Collapse Happens

None of this makes the 300 trillion token ceiling irrelevant — a hard limit on new human text is still a hard limit, and overtraining could bring the reckoning as early as 2027. But the actionable lesson isn’t to panic about the calendar; it’s to ask how a model’s synthetic training data was filtered, and by whom — the same discipline that matters wherever AI systems make decisions with real consequences, like the AI trading agents already losing real money fast. Curation costs compute and headcount, the same scarce resources already strained by the AI data center power bottleneck and the push toward wafer-scale chips built to end the GPU cluster era.

The industry’s own roadmap already assumes synthetic data will supply most future training volume; the only real fork in the road is whether it’s curated like Phi’s textbooks or dumped back in like a feedback loop with no ground truth to correct it. Expect the labs that survive the tightening data supply to publish less about token counts and more about their filtering pipelines, because that’s the actual moat now. Running out of human text was never going to stop AI scaling on its own. Losing the tails while nobody was watching might.

One story like this, most days

Written by a CERN physicist. No spam, unsubscribe any time.

Photon Guy
Photon Guy

Photon Guy writes at the intersection of particle physics and heavy computing infrastructure. He spent years at CERN working on silicon particle detectors — the sensors that catch what the world's largest accelerators smash together — before moving into the data center industry, where he works on the machines that power the internet and AI. ScienceShot is where those two worlds meet: real physics, real engineering, strong opinions, and no press-release rewrites.

Articles: 43

Leave a Reply

Your email address will not be published. Required fields are marked *