The next AI bottleneck may not be compute
For the last few years, the recipe for building better AI has looked fairly straightforward: make the models bigger, give them more computing power, and train them on more data.
That approach has worked remarkably well. But there is a problem hiding in the third part of that recipe.
How much more data is actually available?
It is easy to think of the internet as an endless source of training material. There are billions of webpages, books, articles, images, videos, conversations, and lines of code online. But AI models need useful training data, not just a larger pile of files. As models become more capable, the amount of high-quality data needed to train them grows too.
Research from Google DeepMind made this relationship clearer. Their compute-optimal training work showed that simply making a model larger was not enough: a smaller model trained on significantly more data could outperform a much larger one under a similar compute budget. [1]
Compute only goes so far if you do not have enough good data to use it effectively.
And that creates a different kind of scaling problem.
GPUs can be manufactured. Data centers can be expanded. Training runs can be made larger.
But genuinely new human knowledge does not appear on demand.
The public internet may keep getting bigger, but much of that growth comes from duplicated, outdated, low-quality, restricted, or AI-generated content. More content does not necessarily mean more useful human-generated information.
AI is not literally running out of data today. The problem is that useful human-generated data is finite, and finding enough of it could become a much bigger constraint as models continue to scale.
Scaling has a data bill
Making an AI model bigger sounds like a problem for GPUs. But scaling comes with another bill: data.
As models become larger, they need more training data to make use of that extra capacity. The problem is that high-quality human-generated data does not grow as quickly as the appetite of these models. The internet keeps getting bigger, but that does not mean it is producing an endless supply of useful training material.
Researchers have started putting numbers around this problem. A 2024 study by Villalobos and colleagues estimated that, depending on the assumptions, demand for human-generated training data could begin approaching its available supply somewhere between 2026 and 2032. [2]
| 2024 | 2026 | 2028 | 2030 | 2032 |
That is not a prediction that the internet will suddenly run out of text. It is about whether the supply of usable human-generated data can keep pace with what increasingly large models need.
But raw volume can be misleading. A huge collection of files does not necessarily mean a huge collection of useful information.
So when we talk about the amount of data available to AI, the important number is not everything sitting on the internet.
It is the amount of new, usable information that can actually improve a model.
As training datasets reach billions or even trillions of tokens, finding another huge collection of files is one thing. Finding genuinely useful new information is another.
The problem is not that the internet is becoming small. It is that valuable new data may not be growing fast enough.
So if the internet keeps producing more content, why can’t AI companies simply keep collecting it?
The internet isn’t an infinite dataset
At first, the solution seems obvious: if AI needs more data, just collect more of the internet.
There is certainly a lot to collect. New articles, posts, videos, comments, code, and reviews appear online every day. But more content does not automatically mean more useful training data.
Take duplication and quality. The same information can appear across hundreds of websites, while carefully researched material sits alongside spam, outdated pages, automatically generated text, and plain mistakes.[3] Counting all of it makes the dataset look bigger, but it does not give the model the same amount of new information.Counting all of it makes the dataset look bigger, but it does not give the model the same amount of new information. [4]
There is also the question of what the model has already seen. As datasets get larger, researchers have to check for contamination—for example, when benchmark or evaluation material accidentally ends up in the training data. That can make a model appear better than it really is. [5]
There is also the question of access. Not everything on the internet is available for unrestricted use. Copyright, licensing, privacy, and website terms can all affect whether material can actually be included in a training dataset.
So the amount of content online is very different from the amount that can become useful training data.
The data problem is therefore becoming less about how much content exists and more about how much of it is new, useful, diverse, and actually usable.
And increasingly, some of that new content is being created by AI itself.
Synthetic data has a ceiling
This is where the data problem gets a little more complicated.
If good human-generated data becomes harder to find, the obvious idea is to create more of it ourselves. AI can already generate text, images, code, conversations, and all kinds of other examples at a speed people simply cannot match.
On paper, that sounds almost like an unlimited supply of training data.
And synthetic data can genuinely be useful. It can create situations that are difficult to collect in the real world, generate rare examples, and expand a dataset when gathering real data is expensive or impractical.
But there is a catch.
AI-generated data ultimately comes from what earlier models learned from.
Imagine a model trained mostly on human-created information. It generates millions of pieces of new content, and some of that content ends up being used to train the next model. That model then produces more content, which becomes training material for another model.
Researchers have found that repeatedly training on model-generated data can lead to model collapse, where newer generations lose some of the variety and information present in the original human-generated data. [6]
That does not make synthetic data useless. It can still be valuable for rare situations, controlled examples, edge cases, and tasks where collecting enough real-world data is difficult.
The limitation is simple: generating more examples is not the same as generating more knowledge. A model can produce thousands of variations of something it already understands without adding thousands of genuinely new insights.
Real-world data still brings something synthetic data struggles to reproduce: unpredictability. Human conversations, messy environments, sensor readings, and professional workflows contain details that are difficult to anticipate in advance.
So the future is unlikely to be a choice between human and synthetic data. Real-world data can provide the foundation, while synthetic data can expand it and target specific gaps.
That also changes where AI companies need to look next. The answer may no longer be another corner of the web, but people, specialised industries, software environments, sensors, machines, and the physical world.
the training process. Real-world data is still needed to keep the model grounded in information it could not simply generate from its existing knowledge.
The search for data beyond the web
If the public internet cannot keep supplying useful data forever, AI developers have to start looking elsewhere.
One obvious source is specialised human-generated data. The web may have millions of pages about medicine, engineering, finance, or law, but that does not capture everything professionals actually do. Data collected from real workflows can contain details that are difficult to find through ordinary web scraping.
The same applies to multimodal and real-world data. Audio, video, sensor readings, software interactions, and physical environments can capture information that text alone cannot.
The physical world is another important source. Robotics, for example, needs data about how objects move, how surfaces react, how things go wrong, and how people interact with physical environments. Those details have to be learned from real interactions with the world, not just webpages.
That means collecting data from real interactions with the world.
The same pattern appears in scientific experiments, business operations, and human demonstrations. These sources can show a model how something actually happens, rather than simply describing it in text.
The same pattern appears in scientific experiments, business operations, and human demonstrations. These sources can show a model how something actually happens, rather than simply describing it in text.
That makes these datasets more expensive to build, but also harder to replace.
The next generation of AI may depend less on how much data can be scraped from the internet and more on how much valuable data can be deliberately collected and validated.
And that shift is changing the data pipeline itself.
A new data pipeline is emerging
For a long time, AI training looked fairly simple: collect a lot of data, clean it up, and train the model.
That gets harder when the data comes from specialised industries, real-world interactions, or human demonstrations. You have to be more deliberate about what you collect, who provides it, and whether it is actually useful.
The process starts looking less like scraping and more like a proper pipeline:
The steps are not just cleanup. They decide what information actually reaches the model. And the process does not necessarily stop after training. If evaluation reveals a weakness, that failure can point to a gap in the data, sending the team back to collect and validate more examples. [8]
That creates a feedback loop between data and model performance.
This also makes high-quality data more valuable. A specialised dataset may contain far fewer examples than a massive web scrape, but those examples can be far more useful for a specific task.
The data race is shifting from finding the biggest dataset to building datasets that are new, relevant, diverse, trustworthy, and grounded in the real world.
The bottleneck is shifting
Saying AI is “running out of data” makes it sound like the internet is about to go empty. That is not what is happening. The problem is that the easiest sources of useful human-generated data are becoming harder to scale.
Synthetic data can help, but it cannot fully replace information created through real human experience. The advantage may increasingly come from better, more diverse data that is difficult to collect or reproduce.
That could mean specialised industries, human interactions, scientific work, sensors, machines, and the physical world.
The next data challenge may not be finding the largest dataset. It may be finding data that adds something genuinely new.
AI still has enormous amounts of information left to learn from. But getting to the valuable parts is becoming a more deliberate process. And that could make data collection, curation, and validation just as important to AI development as model architecture and computing power.
In the end, the important question may not be “Who has the most data?”
It may be “Who can find the data that everyone else cannot easily find?”
What this means for the next generation of AI
The data problem does not mean AI has hit a wall. It means the way we get training data has to change.
For years, the internet seemed like an almost endless source of material. If a model needed more examples, there were always more webpages, articles, posts, and code to collect. But that becomes less useful as datasets get bigger and the easiest sources of human-generated content become repetitive, difficult to license, or simply less informative.
The next phase is likely to involve a much wider mix of data.
Human-created web content will still matter. Synthetic data will continue to be useful for filling specific gaps. But AI systems will also need more specialised, real-world data that has been deliberately collected for a purpose.
Huge volume
Create controlled examples
Real interactions
That changes the job of building a training dataset too. It is no longer just about collecting as much as possible. Someone has to figure out what the model actually needs, find the right examples, check whether they are reliable, and make sure the dataset reflects the real world.
This becomes even more important as AI moves beyond text.
A language model can learn from billions of words, but a robot needs information about movement and physical environments. A voice system has to deal with accents, interruptions, background noise, and the way people actually speak. A medical AI needs carefully selected clinical information, not simply another million webpages about medicine.
Different AI systems need different kinds of reality.
That is why the future of AI training may be less about building one enormous dataset from the open internet and more about building many specialised sources of valuable data.
The internet gave AI an enormous head start. But the next wave of progress may depend on information that cannot simply be scraped.
And that could make how data is collected, curated, and validated just as important as the models and hardware used to train AI.
The data problem is becoming a collection problem
AI is not going to suddenly wake up one day with nothing left to learn from.
The problem is more subtle. The internet is still full of content, but finding new, useful human-generated data is getting harder. Adding another million webpages to a dataset does not help much if most of them repeat information a model has already seen.
That means the work is shifting.
Instead of simply collecting more content from the web, AI developers may need to spend more time collecting data directly from people, experts, software environments, machines, and the physical world. They may need to record real interactions, capture unusual situations, and build datasets around specific problems.
Synthetic data will have a role too. It can create extra examples and help cover situations where real data is difficult to collect. But it works best when it is grounded in real information rather than being used as an endless replacement for it.
So the next stage of AI training may look less like “find more data” and more like “find better sources of data.”
The internet gave AI an enormous amount to learn from. What comes next may depend on data that takes more effort to collect, understand, and validate.
And that could make the people and systems building those datasets just as important as the models being trained on them.
References
[1] Training Compute-Optimal Large Language Models
Hoffmann, J., et al. (2022). arXiv:2203.15556.
[2] Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data
Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., & Hobbhahn, M. (2024). Proceedings of Machine Learning Research, 235, 49523– 49544.
[3] The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Gao, L., et al. (2020). arXiv:2101.00027.
[4] Deduplicating Training Data Makes Language Models Better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., & Carlini, N. (2021). arXiv:2107.06499.
[5] Benchmark Data Contamination of Large Language Models: A Survey
Xu, C., Guan, S., Greene, D., & Kechadi, M.-T. (2024). arXiv:2406.04244.
[6] AI Models Collapse When Trained on Recursively Generated Data
Shumailov, I., Shumaylov, Z., Zhao, Y., et al. (2024). Nature, 631, 755–759.
[7] The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Kandpal, N., Lester, B., Raffel, C., et al. (2025). arXiv:2506.05209.
[8] Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., et al. (2024). arXiv:2402.00159.