High-quality text data on the public internet totals approximately 10 trillion tokens. That sounds like a lot. GPT-4 was trained on 6–13 trillion tokens. The models are approaching the wall. Epoch AI warns with 80% certainty that high-quality training data exhausts between 2026 and 2028. Elon Musk agreed in January 2026. Microsoft trained Phi-4 on 400 billion synthetic tokens. NVIDIA acquired Gretel AI for $320 million. OpenAI paid News Corp $250 million over five years. The Anthropic copyright settlement: $1.5 billion — 44% of the entire 2025 training data market. The data famine is real. And the economics it is creating are extraordinary.
Not investment advice. Data sourced from Epoch AI Research 2024, International AI Safety Report 2026 (arXiv), Coherent Market Insights Synthetic Data Market 2026, DataIntelo Dataset Licensing Report 2026, Pebblous AI Licensing Analysis June 2026, Grand View Research, Technavio. All figures current as of June 2026.
The AI revolution has been fuelled by a specific resource: human-generated text. Every book digitised, every Wikipedia article written, every forum post argued, every news article published, every academic paper peer-reviewed — for decades, humanity produced and uploaded text to the internet without knowing it was building the training corpus for the most powerful intelligence amplification systems ever created. By 2024, approximately 10 trillion tokens of high-quality text existed on the public internet — the accumulated written output of an educated civilisation, digitised and made searchable.
GPT-4, one of the most capable language models ever built at the time of its release, was trained on somewhere between 6 and 13 trillion tokens. The models are not approaching the wall. They have hit it. Epoch AI, the most rigorous independent research group tracking AI training data consumption, estimated with 80% certainty that high-quality training data will be exhausted between 2026 and 2028. In January 2026, Elon Musk — who has direct visibility into xAI's data procurement challenges — publicly agreed that "we have exhausted AI training data." Pablo Villalobos, the lead author of the Epoch AI study, confirmed the fundamental scarcity remains real regardless of minor timeline shifts from new techniques.
The overtraining problem makes it worse. AI companies have discovered that "overtraining" — feeding the same data to models multiple times to improve inference efficiency — dramatically accelerates data consumption. Meta's Llama 3 was overtrained by a factor of 10. If labs move toward overtraining factors of 100 — which compute economics encourage — the effective data supply runs out years ahead of the naive timeline. The machine is eating itself out of food while the engineers make it eat faster.
The situation is a gold rush in reverse. In a gold rush, the scarcest resource is capital and labour — the gold is there to be found. In the data famine, the capital and the labour are abundant. The resource itself is running out. And unlike gold, human-generated text cannot be mined faster by throwing more equipment at it. It grows only as fast as humans write — and humans are writing less on the open internet as content moves behind paywalls, into private messages, and onto closed platforms.
The data wall is a text wall. The International AI Safety Report 2026 — the most authoritative recent analysis — makes the hierarchy explicit and the opportunity clear:
The shift from text to image, video, and sensor data is not simply a matter of pointing the training pipeline at different files. A single video frame contains far less semantic information than a paragraph of text. New techniques are required to extract meaningful training signal from visual and temporal data — techniques that are being built right now but are not yet mature. The models that learn to extract semantic signal from video at scale will have access to a data reservoir orders of magnitude larger than the text internet. The race to develop those techniques is the most important frontier in AI research that most people outside AI research are not tracking.
The volume of high-quality public text is finite and being consumed. With overtraining, consumption accelerates. The models that exist today have already trained on most of the high-quality text that exists. Future models cannot simply be trained on "more" — there is no more, unless it is created or synthesised.
The data that does exist — private enterprise data, medical records, legal documents, financial transactions — is behind walls of privacy regulation, competitive sensitivity, and contractual restriction. GDPR, HIPAA, CCPA, and equivalent regulations globally make accessing the most valuable training data legally complex and expensive. The richest data is the least accessible.
AI models trained on copyrighted content without licensing face existential legal liability. The Anthropic copyright settlement — $1.5 billion, 44% of the entire 2025 training data market — demonstrates the scale of the risk. News Corp sued, settled for $250 million over five years. Authors, record labels, image agencies, and news publishers are all in various stages of litigation or negotiation. The cost of legal data is now comparable to the cost of computing.
The most dangerous long-term risk: if AI models train on AI-generated content — their own outputs — errors and biases compound through successive generations, degrading model quality in a phenomenon called "model collapse." A 2024 Nature paper demonstrated this mathematically. As more of the internet fills with AI-generated text, the contamination of future training data becomes an existential quality problem. The internet is already estimated to contain 60%+ AI-generated content in some domains.
The recognition that high-quality human-generated text is scarce has created the fastest-moving licensing market in technology history. In 2024 alone, 91 major AI data licensing deals were disclosed. The pattern: AI companies buying permanent or time-limited rights to train on the content of publishers, news organisations, academic databases, and social media platforms that spent decades building the most valuable text corpus in existence — without knowing it was training data for AI.
If real data is scarce, manufacture it. Synthetic data — artificially generated datasets that mimic the statistical properties of real data — is the most powerful near-term solution to the data famine. Microsoft trained Phi-4, one of its most capable small language models, on 400 billion synthetic tokens with results comparable to models trained on much larger real datasets. Gartner predicted 75% of businesses would use synthetic data for AI training by 2026 — and the prediction looks conservative. The shift from 1% adoption in 2021 to 60% in 2024 is the fastest enterprise technology adoption in recent history.
The market is $0.92 billion in 2026, growing at 34.5% CAGR to $3 billion by 2030. Key players: NVIDIA (acquired Gretel AI for $320 million in March 2025, integrating into Omniverse Replicator with 252 enterprise deployments), MOSTLY AI (structured tabular data for financial services), MDClone (synthetic healthcare data), Synthesis AI, Tonic.ai. The domain where synthetic data works best: verifiable outputs — mathematics, code, formal reasoning — where models can generate solutions and check correctness automatically. The domain where it fails: creative writing, strategic planning, scientific hypothesis generation — wherever there is no objective verifier, errors compound into model collapse.
The shift from free scraping to paid licensing is the most consequential economic restructuring in the content industry since the birth of digital advertising. The global dataset licensing for AI training market was $4.8 billion in 2025 and is projected to reach $22.6 billion by 2034. OpenAI alone has committed 53% of the disclosed deal value — approximately $816 million per year across 34+ deals.
The structural shift is from one-time training dumps to real-time API feeds. Reddit, Wikimedia, and Bloomberg are all moving toward continuous, paid, structured data access rather than historical archives. The most significant trend: content licensing is becoming real-time infrastructure, not historical asset sales. Scale AI established a dominant position with its $650 million DoD contract and its enterprise data labelling business. Appen, Lionbridge, and Sama provide the human annotation layer that makes raw data into training-quality data. The winner-take-most dynamic: $10 million is the minimum disclosed deal size, creating a market where only the largest publishers can participate in the first wave of value capture.
The richest data is behind privacy walls. Federated learning is the technique that unlocks it: instead of moving data to the model, the model moves to the data. A federated learning system trains on data where it lives — in a hospital, a bank, a retailer — without the raw data ever leaving the institution. Only model updates (gradients) are shared, not the underlying records. This makes GDPR, HIPAA, and CCPA compliance structurally embedded in the training process rather than a legal workaround.
Google has pioneered federated learning at scale — Android keyboard predictions are trained on device without user data leaving the phone. Apple uses federated learning across iOS devices. Healthcare is the highest-value application: hospitals with patient records containing 100 million medical imaging studies — among the most valuable training data for medical AI — can participate in model training without violating patient privacy. The economic implication: federated learning creates access to private enterprise data that would cost tens of billions of dollars to license conventionally.
When text runs out, use everything else. The International AI Safety Report 2026 quantified the data hierarchy: image data at 10¹⁴–10¹⁵ tokens, video at 10¹⁵–10¹⁶ tokens, IoT sensor data at 10¹⁷ tokens annually. The shift from text to multimodal training is the single biggest expansion of the addressable data universe available to AI developers — orders of magnitude larger than text. It requires new architectures (vision-language models, video understanding, sensor fusion) but the data itself is abundant.
Tesla's training approach demonstrates this at scale: the fleet of Tesla vehicles generates approximately 1.5 billion miles of real-world driving data annually — a multimodal sensor dataset (cameras, radar, ultrasound, GPS, accelerometers) that cannot be replicated through text or synthetic generation. That physical-world, embodied, sensor data is what trains autonomous driving beyond the capabilities of any text-based model. The companies that build the infrastructure to extract meaningful training signal from video, audio, and sensor data will have access to training data supplies measured in exabytes rather than terabytes.
The most intellectually elegant solution to the data wall is not finding more data — it is getting more out of less. Inference-time scaling is the discovery that letting models "think longer" during use — generating and evaluating multiple reasoning paths before answering — produces dramatically better results without requiring more training data. OpenAI's o1 and o3 models, Anthropic's extended thinking, and Google's Gemini reasoning chains all demonstrate this: a model that reasons through a problem step-by-step outperforms a much larger model that doesn't, using the same training data.
The chain-of-thought breakthrough generates its own training data: models trained on millions of self-generated reasoning chains — where each step can be verified against known answers — create a virtuous cycle of self-improvement that does not require external data. This is the technique behind DeepSeek's efficient training, behind Phi-4's performance despite its small size, and behind the rapid capability improvements of 2025–2026 without proportional increases in training data consumption. Inference-time scaling effectively extends the useful life of existing training data by extracting more signal from it at deployment time.
The long-term winner in the data economy will not be the company that finds the most data — it will be the company that builds the most valuable proprietary data moat. Bloomberg's financial data, Thomson Reuters' legal database, Epic Systems' electronic health records, Palantir's government intelligence data, Salesforce's CRM transaction data — these are datasets that cannot be scraped, purchased, or synthesised. They exist only because the company built the product that generated the data, and the data is only accessible to the company that owns it.
The economic logic: as public data becomes exhausted and licensed data becomes expensive, proprietary data becomes the primary source of competitive differentiation in AI. A model trained on Bloomberg's proprietary financial data will outperform a generic model on financial tasks not because it has more parameters but because it has better data. The implication for investors: companies with large, unique, defensible proprietary datasets are not primarily software companies or AI companies — they are data companies, and their data moats may be the most durable competitive advantages in the AI economy.
Synthetic data solves the quantity problem but not the quality problem. The model collapse risk — AI systems trained on AI-generated content degrading in quality as errors compound — is mathematically real and empirically demonstrated. In domains with verifiable answers (maths, code), synthetic data works brilliantly. In domains without verifiable answers (creative writing, judgement, nuance), synthetic data degrades model quality. As more of the internet fills with AI-generated text, future models face the prospect of training on a corpus increasingly contaminated by the outputs of their predecessors. The 2024 Nature paper demonstrating model collapse is not a hypothetical warning — it is a mathematical proof with empirical confirmation. The data quality problem is harder than the data quantity problem, and solving quantity while ignoring quality is not a solution.
The winner-take-most dynamics in data licensing create a structural inequality that mirrors the content industry's historical pattern. Only the largest publishers — News Corp, Reuters, Shutterstock — have the market power to negotiate nine-figure licensing deals. The $10 million minimum disclosed deal size means that the majority of content creators — independent authors, individual journalists, small publishers, academic researchers — cannot participate in the licensing economy at all. Their work is used for training (or was, before the copyright litigation wave), but they receive no compensation. The Anthropic settlement's $1.5 billion is divided among a class of authors; individual shares will be small. The data licensing market is replicating the pattern of the music streaming revolution: the platforms capture the value, the creators receive fractions.
IoT sensor data is the largest untapped frontier — and also the most geopolitically contested. The 10¹⁷ tokens of annual IoT data includes factory sensors, smart city infrastructure, agricultural monitors, vehicle fleets, and energy grid sensors. Most of this data is owned by corporations or governments in countries with different data sovereignty rules. China's data laws prohibit export of certain categories of data; the EU's data governance framework restricts cross-border flows; India is building a data localisation regime. The geopolitics of data sovereignty are creating the same fragmentation in the data economy that sanctions and export controls are creating in the semiconductor economy — a world of data blocs rather than a global data commons.
Long-horizon thinking on capital, technology, and the forces shaping the next decade of wealth creation. Written from first principles. Not consensus. Not noise.