{"id":936,"date":"2026-10-01T12:03:24","date_gmt":"2026-10-01T12:03:24","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=936"},"modified":"2026-10-01T12:03:24","modified_gmt":"2026-10-01T12:03:24","slug":"ai-is-running-out-of-data-what-happens-when-the-internet-stops-scaling","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/ai-is-running-out-of-data-what-happens-when-the-internet-stops-scaling\/","title":{"rendered":"AI Is Running Out of Data: What Happens When the Internet Stops Scaling?"},"content":{"rendered":"<h2>The next AI bottleneck may not be compute<\/h2>\n<p>For the last few years, the recipe for building better AI has looked fairly straightforward: make the models bigger, give them more computing power, and train them on more data.<\/p>\n<p>That approach has worked remarkably well. But there is a problem hiding in the third part of that recipe.<\/p>\n<p><strong>How much more data is actually available?<\/strong><\/p>\n<p>It is easy to think of the internet as an endless source of training material. There are billions of webpages, books, articles, images, videos, conversations, and lines of code online. But AI models need <strong>useful training data, not just a larger pile of files.<\/strong> As models become more capable, the amount of high-quality data needed to train them grows too.<\/p>\n<p>Research from Google DeepMind made this relationship clearer. Their compute-optimal training work showed that simply making a model larger was not enough: a smaller model trained on significantly more data could outperform a much larger one under a similar compute budget. <a style=\"text-decoration: none;\" href=\"#ref1\">[1]<\/a><\/p>\n<p><strong>Compute only goes so far if you do not have enough good data to use it effectively.<\/strong><\/p>\n<div style=\"max-width: 760px; margin: 24px auto; font-family: Arial, sans-serif;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 8px; flex-wrap: wrap;\">\n<div style=\"background: linear-gradient(135deg, #fffafa, #fcecec); border: 1px solid #f0d2d2; border-radius: 9px; padding: 11px 14px; text-align: center; min-width: 120px;\">\n<div style=\"font-size: 14px; font-weight: bold; color: #793535;\">Larger Models<\/div>\n<div style=\"font-size: 11px; color: #967070; margin-top: 3px;\">More capacity<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"background: linear-gradient(135deg, #fffafa, #fcecec); border: 1px solid #f0d2d2; border-radius: 9px; padding: 11px 14px; text-align: center; min-width: 120px;\">\n<div style=\"font-size: 14px; font-weight: bold; color: #793535;\">More Compute<\/div>\n<div style=\"font-size: 11px; color: #967070; margin-top: 3px;\">More training power<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"background: linear-gradient(135deg, #fffafa, #fcecec); border: 1px solid #f0d2d2; border-radius: 9px; padding: 11px 14px; text-align: center; min-width: 120px;\">\n<div style=\"font-size: 14px; font-weight: bold; color: #793535;\">More Data<\/div>\n<div style=\"font-size: 11px; color: #967070; margin-top: 3px;\">More training examples<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"background: linear-gradient(135deg, #fff8f8, #fbdede); border: 1px solid #e9c3c3; border-radius: 9px; padding: 11px 14px; text-align: center; min-width: 135px;\">\n<div style=\"font-size: 14px; font-weight: bold; color: #713030;\">Data Pressure<\/div>\n<div style=\"font-size: 11px; color: #967070; margin-top: 3px;\">Useful human data is finite<\/div>\n<\/div>\n<\/div>\n<div style=\"text-align: center; margin-top: 10px; font-size: 11px; color: #997070;\">More model scale \u2192 more demand for useful training data<\/div>\n<\/div>\n<p>And that creates a different kind of scaling problem.<\/p>\n<p>GPUs can be manufactured. Data centers can be expanded. Training runs can be made larger.<\/p>\n<p>But <strong>genuinely new human knowledge does not appear on demand.<\/strong><\/p>\n<p>The public internet may keep getting bigger, but much of that growth comes from duplicated, outdated, low-quality, restricted, or AI-generated content. More content does not necessarily mean more <strong>useful human-generated information<\/strong>.<\/p>\n<p>AI is not literally running out of data today. The problem is that <strong>useful human-generated data is finite<\/strong>, and finding enough of it could become a much bigger constraint as models continue to scale.<\/p>\n<div>\n<h2>Scaling has a data bill<\/h2>\n<p>Making an AI model bigger sounds like a problem for GPUs. But scaling comes with another bill: <strong>data<\/strong>.<\/p>\n<p>As models become larger, they need more training data to make use of that extra capacity. The problem is that high-quality human-generated data does not grow as quickly as the appetite of these models. The internet keeps getting bigger, but that does not mean it is producing an endless supply of useful training material.<\/p>\n<p>Researchers have started putting numbers around this problem. A 2024 study by Villalobos and colleagues estimated that, depending on the assumptions, demand for human-generated training data could begin approaching its available supply somewhere between <strong>2026 and 2032<\/strong>. <a style=\"text-decoration: none;\" href=\"#ref2\">[2]<\/a><\/p>\n<div style=\"max-width: 680px; margin: 28px auto; font-family: Arial,sans-serif;\">\n<div style=\"text-align: center; font-size: 14px; font-weight: bold; color: #743333; margin-bottom: 18px;\">When training-data demand starts catching up<\/div>\n<p><!-- AVAILABLE HUMAN DATA --><\/p>\n<div style=\"margin-bottom: 20px;\">\n<div style=\"font-size: 11px; font-weight: bold; color: #823939; margin-bottom: 7px;\">Available human-generated data<\/div>\n<div style=\"width: 100%; height: 14px; background: linear-gradient(90deg,#f9dddd,#efc1c1); border-radius: 8px;\"><\/div>\n<div style=\"font-size: 10px; color: #987070; margin-top: 5px;\">Large, but ultimately finite supply<\/div>\n<\/div>\n<p><!-- TRAINING DATA DEMAND --><\/p>\n<div style=\"margin-bottom: 18px;\">\n<div style=\"font-size: 11px; font-weight: bold; color: #9a4242; margin-bottom: 7px;\">Training-data demand<\/div>\n<div style=\"width: 100%; height: 14px; border-radius: 8px; overflow: hidden; background: #f7dddd;\">\n<div style=\"width: 40%; height: 14px; display: inline-block; vertical-align: top; background: #f2cccc;\"><\/div>\n<div style=\"width: 28%; height: 14px; display: inline-block; vertical-align: top; background: #df9999;\"><\/div>\n<div style=\"width: 32%; height: 14px; display: inline-block; vertical-align: top; background: #bd5959;\"><\/div>\n<\/div>\n<div style=\"font-size: 10px; color: #987070; margin-top: 5px;\">Growing as models require more training data<\/div>\n<\/div>\n<p><!-- TIMELINE --><\/p>\n<div style=\"width: 100%; margin-top: 18px;\">\n<p><!-- YEARS --><\/p>\n<table style=\"width: 100%; border-collapse: collapse; table-layout: fixed; margin: 0; padding: 0; border: none;\">\n<tbody>\n<tr>\n<td style=\"width: 20%; text-align: left; padding: 0; border: none; font-size: 10px; color: #987070;\">2024<\/td>\n<td style=\"width: 20%; text-align: center; padding: 0; border: none; font-size: 10px; color: #987070;\">2026<\/td>\n<td style=\"width: 20%; text-align: center; padding: 0; border: none; font-size: 10px; color: #987070;\">2028<\/td>\n<td style=\"width: 20%; text-align: center; padding: 0; border: none; font-size: 10px; color: #987070;\">2030<\/td>\n<td style=\"width: 20%; text-align: right; padding: 0; border: none; font-size: 10px; color: #987070;\">2032<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><!-- TIMELINE BAR --><\/p>\n<div style=\"width: 100%; height: 5px; margin-top: 7px; border-radius: 4px; background: linear-gradient( 90deg, #f7e5e5 0%, #f7e5e5 40%, #e9aaaa 40%, #c96a6a 100% );\"><\/div>\n<p><!-- PRESSURE LABEL --><\/p>\n<div style=\"text-align: center; margin-top: 8px; font-size: 10px; font-weight: bold; color: #873838;\">Projected pressure window: 2026\u20132032<\/div>\n<\/div>\n<p><!-- CAPTION --><\/p>\n<div style=\"max-width: 600px; margin: 14px auto 0; text-align: center; font-size: 10.5px; line-height: 1.5; color: #8d6868;\">Conceptual view based on the study&#8217;s projections. The exact timing depends on the assumptions used.<\/div>\n<\/div>\n<p>That is not a prediction that the internet will suddenly run out of text. It is about whether the supply of usable human-generated data can keep pace with what increasingly large models need.<\/p>\n<p>But raw volume can be misleading. A huge collection of files does not necessarily mean a huge collection of useful information.<br \/>\nSo when we talk about the amount of data available to AI, the important number is not everything sitting on the internet.<\/p>\n<p>It is the amount of <strong>new, usable information that can actually improve a model.<\/strong><\/p>\n<p>As training datasets reach billions or even trillions of tokens, finding another huge collection of files is one thing. Finding genuinely useful new information is another.<\/p>\n<p><strong>The problem is not that the internet is becoming small. It is that valuable new data may not be growing fast enough.<\/strong><\/p>\n<p>So if the internet keeps producing more content, why can&#8217;t AI companies simply keep collecting it?<\/p>\n<\/div>\n<div>\n<h2>The internet isn&#8217;t an infinite dataset<\/h2>\n<p>At first, the solution seems obvious: if AI needs more data, just collect more of the internet.<\/p>\n<p>There is certainly a lot to collect. New articles, posts, videos, comments, code, and reviews appear online every day. But <strong>more content does not automatically mean more useful training data.<\/strong><\/p>\n<div style=\"width: 100%; margin: 28px 0; padding: 2px 0 2px 16px; border-left: 3px solid #d98f8f; box-sizing: border-box; font-family: Arial,sans-serif;\">\n<div style=\"font-size: 15px; font-weight: bold; line-height: 1.5; color: #743333; margin-bottom: 6px;\">More content does not mean more useful data.<\/div>\n<div style=\"font-size: 14px; line-height: 1.7; color: #555;\">The internet can keep getting bigger without providing the same amount of genuinely new information. A large dataset may contain repeated pages, outdated material, low-quality content, restricted sources, or information that has already appeared in previous training datasets. What matters is not simply how much content exists, but how much of it adds something new and useful to the model.<\/div>\n<\/div>\n<p>Take duplication and quality. The same information can appear across hundreds of websites, while carefully researched material sits alongside spam, outdated pages, automatically generated text, and plain mistakes.<a style=\"text-decoration: none;\" href=\"#ref3\">[3]<\/a> Counting all of it makes the dataset look bigger, but it does not give the model the same amount of new information.Counting all of it makes the dataset look bigger, but it does not give the model the same amount of new information. <a style=\"text-decoration: none;\" href=\"#ref4\">[4]<\/a><\/p>\n<p>There is also the question of what the model has already seen. As datasets get larger, researchers have to check for <strong>contamination<\/strong>\u2014for example, when benchmark or evaluation material accidentally ends up in the training data. That can make a model appear better than it really is. <a style=\"text-decoration: none;\" href=\"#ref5\">[5]<\/a><\/p>\n<p>There is also the question of access. <strong>Not everything on the internet is available for unrestricted use.<\/strong> Copyright, licensing, privacy, and website terms can all affect whether material can actually be included in a training dataset.<\/p>\n<p>So the amount of content online is very different from the amount that can become useful training data.<\/p>\n<p>The data problem is therefore becoming less about <strong>how much content exists<\/strong> and more about how much of it is new, useful, diverse, and actually usable.<\/p>\n<p>And increasingly, some of that new content is being created by <strong>AI itself.<\/strong><\/p>\n<div>\n<h2>Synthetic data has a ceiling<\/h2>\n<p>This is where the data problem gets a little more complicated.<\/p>\n<p>If good human-generated data becomes harder to find, the obvious idea is to create more of it ourselves. AI can already generate text, images, code, conversations, and all kinds of other examples at a speed people simply cannot match.<br \/>\nOn paper, that sounds almost like an unlimited supply of training data.<\/p>\n<p>And synthetic data can genuinely be useful. It can create situations that are difficult to collect in the real world, generate rare examples, and expand a dataset when gathering real data is expensive or impractical.<\/p>\n<p>But there is a catch.<\/p>\n<p><strong>AI-generated data ultimately comes from what earlier models learned from.<\/strong><\/p>\n<p>Imagine a model trained mostly on human-created information. It generates millions of pieces of new content, and some of that content ends up being used to train the next model. That model then produces more content, which becomes training material for another model.<\/p>\n<div style=\"width: 100%; margin: 28px 0; font-family: Arial,sans-serif;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 10px; flex-wrap: wrap;\">\n<div style=\"flex: 1; min-width: 130px; padding: 12px 10px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Human Data<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Original information<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 130px; padding: 12px 10px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">AI Model<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Learns from the data<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 130px; padding: 12px 10px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Synthetic Data<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Model-generated content<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 130px; padding: 12px 10px; text-align: center; background: linear-gradient(135deg,#fff8f8,#fbdede); border: 1px solid #e8c4c4; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #713030;\">Next Model<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Trained on the mixture<\/div>\n<\/div>\n<\/div>\n<div style=\"text-align: center; margin-top: 10px; font-size: 11px; color: #966b6b;\">Repeating the cycle can reduce the amount of fresh human information entering training.<\/div>\n<div>\n<p>Researchers have found that repeatedly training on model-generated data can lead to <strong>model collapse<\/strong>, where newer generations lose some of the variety and information present in the original human-generated data. <a style=\"text-decoration: none;\" href=\"#ref6\">[6]<\/a><\/p>\n<p>That does not make synthetic data useless. It can still be valuable for rare situations, controlled examples, edge cases, and tasks where collecting enough real-world data is difficult.<\/p>\n<p>The limitation is simple: <strong>generating more examples is not the same as generating more knowledge.<\/strong> A model can produce thousands of variations of something it already understands without adding thousands of genuinely new insights.<\/p>\n<p>Real-world data still brings something synthetic data struggles to reproduce: unpredictability. Human conversations, messy environments, sensor readings, and professional workflows contain details that are difficult to anticipate in advance.<\/p>\n<p>So the future is unlikely to be a choice between human and synthetic data. Real-world data can provide the foundation, while synthetic data can expand it and target specific gaps.<\/p>\n<p>That also changes where AI companies need to look next. The answer may no longer be another corner of the web, but <strong>people, specialised industries, software environments, sensors, machines, and the physical world.<\/strong><\/p>\n<\/div>\n<\/div>\n<\/div>\n<div>\n<div>\n<div style=\"width: 100%; margin: 28px 0; padding: 2px 0 2px 16px; border-left: 3px solid #d98f8f; box-sizing: border-box; font-family: Arial,sans-serif;\">\n<div style=\"font-size: 15px; font-weight: bold; line-height: 1.5; color: #743333; margin-bottom: 6px;\">Synthetic data can expand knowledge, but it cannot endlessly create new knowledge.<\/div>\n<div style=\"font-size: 14px; line-height: 1.7; color: #555;\">A model can generate thousands of new examples from patterns it has already learned. Those examples can be useful for practice, edge cases, and filling specific gaps. But generating more variations does not automatically introduce new information into<br \/>\nthe training process. Real-world data is still needed to keep the model grounded in information it could not simply generate from its existing knowledge.<\/div>\n<\/div>\n<\/div>\n<div>\n<h2>The search for data beyond the web<\/h2>\n<p>If the public internet cannot keep supplying useful data forever, AI developers have to start looking elsewhere.<\/p>\n<p>One obvious source is <strong>specialised human-generated data<\/strong>. The web may have millions of pages about medicine, engineering, finance, or law, but that does not capture everything professionals actually do. Data collected from real workflows can contain details that are difficult to find through ordinary web scraping.<\/p>\n<p>The same applies to multimodal and real-world data. Audio, video, sensor readings, software interactions, and physical environments can capture information that text alone cannot.<\/p>\n<div style=\"width: 100%; margin: 28px 0; font-family: Arial,sans-serif; overflow-x: auto;\">\n<table style=\"width: 100%; border-collapse: collapse; font-size: 13px; line-height: 1.5; border: 1px solid #efcfcf; border-radius: 9px; overflow: hidden;\">\n<thead>\n<tr style=\"background: linear-gradient(135deg,#fffafa,#fcecec);\">\n<th style=\"padding: 11px 12px; text-align: left; color: #743333; border-bottom: 1px solid #efcfcf;\">Data source<\/th>\n<th style=\"padding: 11px 12px; text-align: left; color: #743333; border-bottom: 1px solid #efcfcf;\">What it adds<\/th>\n<th style=\"padding: 11px 12px; text-align: left; color: #743333; border-bottom: 1px solid #efcfcf;\">Main challenge<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"padding: 11px 12px; color: #783434; font-weight: bold; border-bottom: 1px solid #f2dddd;\">Public web<\/td>\n<td style=\"padding: 11px 12px; color: #555; border-bottom: 1px solid #f2dddd;\">Huge volume and broad coverage<\/td>\n<td style=\"padding: 11px 12px; color: #666; border-bottom: 1px solid #f2dddd;\">Duplication, quality, licensing and contamination<\/td>\n<\/tr>\n<tr style=\"background: #fffafa;\">\n<td style=\"padding: 11px 12px; color: #783434; font-weight: bold; border-bottom: 1px solid #f2dddd;\">Human &amp; domain data<\/td>\n<td style=\"padding: 11px 12px; color: #555; border-bottom: 1px solid #f2dddd;\">Expert knowledge and real workflows<\/td>\n<td style=\"padding: 11px 12px; color: #666; border-bottom: 1px solid #f2dddd;\">Expensive and harder to collect<\/td>\n<\/tr>\n<tr>\n<td style=\"padding: 11px 12px; color: #783434; font-weight: bold; border-bottom: 1px solid #f2dddd;\">Synthetic data<\/td>\n<td style=\"padding: 11px 12px; color: #555; border-bottom: 1px solid #f2dddd;\">Controlled examples and rare edge cases<\/td>\n<td style=\"padding: 11px 12px; color: #666; border-bottom: 1px solid #f2dddd;\">May reproduce existing model patterns<\/td>\n<\/tr>\n<tr style=\"background: #fffafa;\">\n<td style=\"padding: 11px 12px; color: #783434; font-weight: bold;\">Real-world multimodal data<\/td>\n<td style=\"padding: 11px 12px; color: #555;\">Audio, video, sensors and physical interactions<\/td>\n<td style=\"padding: 11px 12px; color: #666;\">Collection, annotation and validation are complex<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<div style=\"text-align: center; margin-top: 9px; font-size: 10.5px; color: #987070;\">The next data advantage may come from harder-to-collect sources rather than larger web-scale datasets. <a style=\"text-decoration: none;\" href=\"#ref7\">[7]<\/a><\/div>\n<\/div>\n<p>The physical world is another important source. Robotics, for example, needs data about how objects move, how surfaces react, how things go wrong, and how people interact with physical environments. Those details have to be learned from <strong>real interactions with the world<\/strong>, not just webpages.<\/p>\n<p>That means collecting data from <strong>real interactions with the world<\/strong>.<\/p>\n<p>The same pattern appears in scientific experiments, business operations, and human demonstrations. These sources can show a model <strong>how something actually happens<\/strong>, rather than simply describing it in text.<\/p>\n<p>The same pattern appears in scientific experiments, business operations, and human demonstrations. These sources can show a model <strong>how something actually happens<\/strong>, rather than simply describing it in text.<\/p>\n<p>That makes these datasets more expensive to build, but also harder to replace.<\/p>\n<p><strong>The next generation of AI may depend less on how much data can be scraped from the internet and more on how much valuable data can be deliberately collected and validated.<\/strong><\/p>\n<p>And that shift is changing the data pipeline itself.<\/p>\n<div>\n<h2>A new data pipeline is emerging<\/h2>\n<p>For a long time, AI training looked fairly simple: <strong>collect a lot of data, clean it up, and train the model.<\/strong><\/p>\n<p>That gets harder when the data comes from specialised industries, real-world interactions, or human demonstrations. You have to be more deliberate about <strong>what you collect, who provides it, and whether it is actually useful.<\/strong><\/p>\n<p>The process starts looking less like scraping and more like a proper pipeline:<\/p>\n<div style=\"width: 100%; margin: 30px 0; font-family: Arial,sans-serif;\">\n<p><!-- Top row --><\/p>\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 8px; flex-wrap: wrap;\">\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Collect<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Gather raw data<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Filter<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Remove noise<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Curate<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Select useful examples<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Annotate<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Add context<\/div>\n<\/div>\n<\/div>\n<p><!-- Middle transition --><\/p>\n<div style=\"display: flex; justify-content: flex-end; padding: 6px 12%; font-size: 20px; color: #b96a6a;\">\u2193<\/div>\n<p><!-- Bottom row --><\/p>\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 8px; flex-wrap: wrap;\">\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fff8f8,#fbdede); border: 1px solid #e8c4c4; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #713030;\">Validate<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Check quality<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2190<\/div>\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Train<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Build the model<\/div>\n<\/div>\n<div style=\"font-size: 19px; color: #b96a6a;\">\u2190<\/div>\n<div style=\"flex: 1; min-width: 115px; padding: 12px 9px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Evaluate<\/div>\n<div style=\"font-size: 10px; color: #967070; margin-top: 4px;\">Find remaining gaps<\/div>\n<\/div>\n<\/div>\n<p><!-- Return arrow --><\/p>\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 10px; margin-top: 10px;\">\n<div style=\"font-size: 20px; color: #b96a6a;\"><\/div>\n<div style=\"font-size: 11px; font-weight: 600; color: #8f6464;\">Feedback from evaluation drives the next round of data collection<\/div>\n<\/div>\n<\/div>\n<p>The steps are not just cleanup. They decide what information actually reaches the model. And the process does not necessarily stop after training. If evaluation reveals a weakness, that failure can point to a gap in the data, sending the team back to collect and validate more examples. <a style=\"text-decoration: none;\" href=\"#ref8\">[8]<\/a><\/p>\n<p>That creates a feedback loop between <strong>data and model performance<\/strong>.<\/p>\n<p>This also makes high-quality data more valuable. A specialised dataset may contain far fewer examples than a massive web scrape, but those examples can be far more useful for a specific task.<\/p>\n<p>The data race is shifting from finding the biggest dataset to building datasets that are <strong>new, relevant, diverse, trustworthy, and grounded in the real world.<\/strong><\/p>\n<h2>The bottleneck is shifting<\/h2>\n<p>Saying AI is \u201crunning out of data\u201d makes it sound like the internet is about to go empty. That is not what is happening. The problem is that the easiest sources of useful human-generated data are becoming harder to scale.<\/p>\n<p>Synthetic data can help, but it cannot fully replace information created through real human experience. The advantage may increasingly come from <strong>better, more diverse data that is difficult to collect or reproduce.<\/strong><\/p>\n<p>That could mean specialised industries, human interactions, scientific work, sensors, machines, and the physical world.<\/p>\n<p>The next data challenge may not be finding the largest dataset. It may be finding <strong>data that adds something genuinely new.<\/strong><\/p>\n<p>AI still has enormous amounts of information left to learn from. But getting to the valuable parts is becoming a more deliberate process. And that could make data collection, curation, and validation just as important to AI development as model architecture and computing power.<\/p>\n<p>In the end, the important question may not be <strong>\u201cWho has the most data?\u201d<\/strong><\/p>\n<p>It may be <strong>\u201cWho can find the data that everyone else cannot easily find?\u201d<\/strong><\/p>\n<div class=\"group\/user-message flex flex-col items-end gap-2\" data-chatgpt-search-unit-key=\"fallback-turn-28:0:user\" data-chatgpt-search-message-ids=\"948e75bf-221a-48db-a0ab-83c8b29609e1\">\n<div class=\"w-full\" data-content-search-unit-key=\"fallback-turn-28:0:user\">\n<div class=\"group\/user-message flex w-full flex-col gap-1 items-end\">\n<div class=\"bg-user-message text-user-message min-w-0 max-w-(--user-chat-width) overflow-hidden break-words px-(--thread-content-margin) [&amp;_.contain-inline-size]:[contain:initial] bubble-vBnidZ py-2.5 rounded-2xl relative text-start\" data-user-message-bubble=\"true\">\n<div class=\"flex flex-col items-end gap-1\">\n<div class=\"relative w-full min-w-0 text-size-chat leading-relaxed\">\n<div class=\"\" data-search-result-target=\"\">\n<div>\n<div class=\"text-size-chat whitespace-pre-wrap\" dir=\"auto\">humanise it<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"block-BQZwFn\">\n<div data-content-search-unit-key=\"fallback-turn-28:1:assistant\" data-chatgpt-search-unit-key=\"fallback-turn-28:1:assistant\" data-chatgpt-search-message-ids=\"2fae93ab-56a3-4aff-82bf-6053d3118bef 2fae93ab-56a3-4aff-82bf-6053d3118bef\">\n<div class=\"group flex min-w-0 flex-col\" data-chatgpt-selection-conversation-id=\"local-chatgpt:dd0411e9-ed89-48bc-95b9-2671ffefd8d8\" data-chatgpt-selection-message-id=\"2fae93ab-56a3-4aff-82bf-6053d3118bef\">\n<div class=\"MarkdownRoot-rZKhxa [&amp;&gt;*:first-child]:mt-0 [&amp;&gt;*:last-child]:mb-0 [&amp;&gt;*:last-child]:mb-0 [&amp;&gt;ol:first-child]:mt-0 [&amp;&gt;ul:first-child]:mt-0\" dir=\"auto\" data-selected-text-overlay-target=\"_r_2ro_\" data-markdown-text-style=\"assistant-message\">\n<h2 class=\"Heading-Ro3lZr\" dir=\"auto\">What this means for the next generation of AI<\/h2>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">The data problem does not mean AI has hit a wall. It means <strong class=\"font-semibold\">the way we get training data has to change.<\/strong><\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">For years, the internet seemed like an almost endless source of material. If a model needed more examples, there were always more webpages, articles, posts, and code to collect. But that becomes less useful as datasets get bigger and the easiest sources of human-generated content become repetitive, difficult to license, or simply less informative.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">The next phase is likely to involve a much wider mix of data.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">Human-created web content will still matter. Synthetic data will continue to be useful for filling specific gaps. But AI systems will also need more <strong class=\"font-semibold\">specialised, real-world data that has been deliberately collected for a purpose.<\/strong><\/p>\n<div style=\"width: 100%; margin: 28px 0; font-family: Arial,sans-serif;\">\n<div style=\"display: flex; align-items: stretch; justify-content: center; gap: 10px; flex-wrap: wrap;\">\n<div style=\"flex: 1; min-width: 170px; padding: 14px 12px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Web Data<\/div>\n<div style=\"font-size: 10.5px; line-height: 1.5; color: #8f6b6b; margin-top: 5px;\">Broad coverage<br \/>\nHuge volume<\/div>\n<\/div>\n<div style=\"display: flex; align-items: center; font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 170px; padding: 14px 12px; text-align: center; background: linear-gradient(135deg,#fffafa,#fcecec); border: 1px solid #efcfcf; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #783434;\">Synthetic Data<\/div>\n<div style=\"font-size: 10.5px; line-height: 1.5; color: #8f6b6b; margin-top: 5px;\">Fill specific gaps<br \/>\nCreate controlled examples<\/div>\n<\/div>\n<div style=\"display: flex; align-items: center; font-size: 19px; color: #b96a6a;\">\u2192<\/div>\n<div style=\"flex: 1; min-width: 170px; padding: 14px 12px; text-align: center; background: linear-gradient(135deg,#fff8f8,#fbdede); border: 1px solid #e8c4c4; border-radius: 9px; box-sizing: border-box;\">\n<div style=\"font-size: 13px; font-weight: bold; color: #713030;\">Real-World Data<\/div>\n<div style=\"font-size: 10.5px; line-height: 1.5; color: #8f6b6b; margin-top: 5px;\">Specialised knowledge<br \/>\nReal interactions<\/div>\n<\/div>\n<\/div>\n<div style=\"text-align: center; margin-top: 11px; font-size: 10.5px; color: #987070;\">The next generation of AI training will use a wider mix of data sources.<\/div>\n<\/div>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">That changes the job of building a training dataset too. It is no longer just about collecting as much as possible. Someone has to figure out what the model actually needs, find the right examples, check whether they are reliable, and make sure the dataset reflects the real world.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">This becomes even more important as AI moves beyond text.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">A language model can learn from billions of words, but a robot needs information about movement and physical environments. A voice system has to deal with accents, interruptions, background noise, and the way people actually speak. A medical AI needs carefully selected clinical information, not simply another million webpages about medicine.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\"><strong class=\"font-semibold\">Different AI systems need different kinds of reality.<\/strong><\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">That is why the future of AI training may be less about building one enormous dataset from the open internet and more about building many specialised sources of valuable data.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">The internet gave AI an enormous head start. But the next wave of progress may depend on information that cannot simply be scraped.<\/p>\n<p class=\"Paragraph-kKnbIo\" dir=\"auto\">And that could make <strong class=\"font-semibold\">how data is collected, curated, and validated<\/strong> just as important as the models and hardware used to train AI.<\/p>\n<div>\n<h2>The data problem is becoming a collection problem<\/h2>\n<p>AI is not going to suddenly wake up one day with nothing left to learn from.<\/p>\n<p>The problem is more subtle. The internet is still full of content, but finding <strong>new, useful human-generated data<\/strong> is getting harder. Adding another million webpages to a dataset does not help much if most of them repeat information a model has already seen.<\/p>\n<p>That means the work is shifting.<\/p>\n<p>Instead of simply collecting more content from the web, AI developers may need to spend more time collecting data directly from people, experts, software environments, machines, and the physical world. They may need to record real interactions, capture unusual situations, and build datasets around specific problems.<\/p>\n<p>Synthetic data will have a role too. It can create extra examples and help cover situations where real data is difficult to collect. But it works best when it is grounded in real information rather than being used as an endless replacement for it.<\/p>\n<p>So the next stage of AI training may look less like <strong>\u201cfind more data\u201d<\/strong> and more like <strong>\u201cfind better sources of data.\u201d<\/strong><\/p>\n<p>The internet gave AI an enormous amount to learn from. What comes next may depend on data that takes more effort to collect, understand, and validate.<\/p>\n<p>And that could make the people and systems building those datasets just as important as the models being trained on them.<\/p>\n<div style=\"font-family: Arial,sans-serif; font-size: 13px; line-height: 1.65; color: #555;\">\n<h2>References<\/h2>\n<div style=\"font-family: Arial,sans-serif; font-size: 13px; line-height: 1.6; color: #555;\">\n<p id=\"ref1\" style=\"margin: 0 0 14px 0;\"><strong>[1]<\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2203.15556\" target=\"_blank\" rel=\"noopener\"><em> Training Compute-Optimal Large Language Models<\/em><\/a><br \/>\nHoffmann, J., et al. (2022). arXiv:2203.15556.<\/p>\n<p id=\"ref2\" style=\"margin: 0 0 14px 0;\"><strong>[2] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/proceedings.mlr.press\/v235\/villalobos24a.html\" target=\"_blank\" rel=\"noopener\"><em>Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data<\/em><\/a><br \/>\nVillalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., &amp; Hobbhahn, M. (2024). Proceedings of Machine Learning Research, 235, 49523\u2013\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 49544.<\/p>\n<p id=\"ref3\" style=\"margin: 0 0 14px 0;\"><strong>[3] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2101.00027\" target=\"_blank\" rel=\"noopener\"><em>The Pile: An 800GB Dataset of Diverse Text for Language Modeling<\/em><\/a><br \/>\nGao, L., et al. (2020). arXiv:2101.00027.<\/p>\n<p id=\"ref4\" style=\"margin: 0 0 14px 0;\"><strong>[4] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2107.06499\" target=\"_blank\" rel=\"noopener\"><em>Deduplicating Training Data Makes Language Models Better<\/em><\/a><br \/>\nLee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., &amp; Carlini, N. (2021). arXiv:2107.06499.<\/p>\n<p id=\"ref5\" style=\"margin: 0 0 14px 0;\"><strong>[5] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2406.04244\" target=\"_blank\" rel=\"noopener\"><em>Benchmark Data Contamination of Large Language Models: A Survey<\/em><\/a><br \/>\nXu, C., Guan, S., Greene, D., &amp; Kechadi, M.-T. (2024). arXiv:2406.04244.<\/p>\n<p id=\"ref6\" style=\"margin: 0 0 14px 0;\"><strong>[6] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/www.nature.com\/articles\/s41586-024-07566-y\" target=\"_blank\" rel=\"noopener\"><em>AI Models Collapse When Trained on Recursively Generated Data<\/em><\/a><br \/>\nShumailov, I., Shumaylov, Z., Zhao, Y., et al. (2024). <em>Nature<\/em>, 631, 755\u2013759.<\/p>\n<p id=\"ref7\" style=\"margin: 0 0 14px 0;\"><strong>[7] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2506.05209\" target=\"_blank\" rel=\"noopener\"><em>The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text<\/em><\/a><br \/>\nKandpal, N., Lester, B., Raffel, C., et al. (2025). arXiv:2506.05209.<\/p>\n<p id=\"ref8\" style=\"margin: 0;\"><strong>[8] <\/strong><a style=\"color: #783434; text-decoration: none;\" href=\"https:\/\/arxiv.org\/abs\/2402.00159\" target=\"_blank\" rel=\"noopener\"><em>Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research<\/em><\/a><br \/>\nSoldaini, L., Kinney, R., Bhagia, A., et al. (2024). arXiv:2402.00159.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>The next AI bottleneck may not be compute For the last few years, the recipe for building better AI has looked fairly straightforward: make the models bigger, give them more\u2026<\/p>\n","protected":false},"author":1,"featured_media":940,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22,3],"tags":[],"class_list":["post-936","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","category-tech-update"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/936","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=936"}],"version-history":[{"count":4,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/936\/revisions"}],"predecessor-version":[{"id":941,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/936\/revisions\/941"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/940"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=936"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=936"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=936"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}