NGE · Investment Letter · Issue 101 · June 2026 · Data Economics · AI Infrastructure

The Data
Famine:
Economics of Scarcity,
Synthetic Solutions,
and the Race to
Feed the AI.

High-quality text data on the public internet totals approximately 10 trillion tokens. That sounds like a lot. GPT-4 was trained on 6–13 trillion tokens. The models are approaching the wall. Epoch AI warns with 80% certainty that high-quality training data exhausts between 2026 and 2028. Elon Musk agreed in January 2026. Microsoft trained Phi-4 on 400 billion synthetic tokens. NVIDIA acquired Gretel AI for $320 million. OpenAI paid News Corp $250 million over five years. The Anthropic copyright settlement: $1.5 billion — 44% of the entire 2025 training data market. The data famine is real. And the economics it is creating are extraordinary.

Not investment advice. Data sourced from Epoch AI Research 2024, International AI Safety Report 2026 (arXiv), Coherent Market Insights Synthetic Data Market 2026, DataIntelo Dataset Licensing Report 2026, Pebblous AI Licensing Analysis June 2026, Grand View Research, Technavio. All figures current as of June 2026.

The Wall — What the Numbers Actually Mean

Every model that ever trained
on the internet
was eating from
the same finite plate.

The AI revolution has been fuelled by a specific resource: human-generated text. Every book digitised, every Wikipedia article written, every forum post argued, every news article published, every academic paper peer-reviewed — for decades, humanity produced and uploaded text to the internet without knowing it was building the training corpus for the most powerful intelligence amplification systems ever created. By 2024, approximately 10 trillion tokens of high-quality text existed on the public internet — the accumulated written output of an educated civilisation, digitised and made searchable.

GPT-4, one of the most capable language models ever built at the time of its release, was trained on somewhere between 6 and 13 trillion tokens. The models are not approaching the wall. They have hit it. Epoch AI, the most rigorous independent research group tracking AI training data consumption, estimated with 80% certainty that high-quality training data will be exhausted between 2026 and 2028. In January 2026, Elon Musk — who has direct visibility into xAI's data procurement challenges — publicly agreed that "we have exhausted AI training data." Pablo Villalobos, the lead author of the Epoch AI study, confirmed the fundamental scarcity remains real regardless of minor timeline shifts from new techniques.

The overtraining problem makes it worse. AI companies have discovered that "overtraining" — feeding the same data to models multiple times to improve inference efficiency — dramatically accelerates data consumption. Meta's Llama 3 was overtrained by a factor of 10. If labs move toward overtraining factors of 100 — which compute economics encourage — the effective data supply runs out years ahead of the naive timeline. The machine is eating itself out of food while the engineers make it eat faster.

10T
High-quality text tokens on the public internet — the ceiling GPT-4 nearly reached alone
10×
Overtraining factor of Meta's Llama 3 — effectively multiplying data consumption
$4.8B
Global dataset licensing for AI training market in 2025 — growing to $22.6B by 2034

The situation is a gold rush in reverse. In a gold rush, the scarcest resource is capital and labour — the gold is there to be found. In the data famine, the capital and the labour are abundant. The resource itself is running out. And unlike gold, human-generated text cannot be mined faster by throwing more equipment at it. It grows only as fast as humans write — and humans are writing less on the open internet as content moves behind paywalls, into private messages, and onto closed platforms.

The Data Hierarchy — Not All Data Is Equal

Text is nearly exhausted.
But the data universe
is vastly larger
than text alone.

The data wall is a text wall. The International AI Safety Report 2026 — the most authoritative recent analysis — makes the hierarchy explicit and the opportunity clear:

High-quality text
Books, papers, curated web
~10¹³ tokens
⚠️ Nearly exhausted · Models approaching ceiling
Image data
Photos, illustrations, art
10¹⁴–10¹⁵
Larger reserve · Late 2030s–2040s exhaustion · Challenge: less semantic richness per token
Video data
YouTube, streaming, film
10¹⁵–10¹⁶
Enormous reserve · 2030–2060 range · Requires new extraction techniques for meaning
IoT sensor data
Devices, machines, environments
10¹⁷/year
Effectively infinite annual generation · The frontier of AI training data · Domain-specific and physical-world grounded

The shift from text to image, video, and sensor data is not simply a matter of pointing the training pipeline at different files. A single video frame contains far less semantic information than a paragraph of text. New techniques are required to extract meaningful training signal from visual and temporal data — techniques that are being built right now but are not yet mature. The models that learn to extract semantic signal from video at scale will have access to a data reservoir orders of magnitude larger than the text internet. The race to develop those techniques is the most important frontier in AI research that most people outside AI research are not tracking.

The Four Crises Within the Famine

Not one problem.
Four simultaneous crises
each requiring
a different solution.

📉
Crisis 1: Quantitative Exhaustion

The volume of high-quality public text is finite and being consumed. With overtraining, consumption accelerates. The models that exist today have already trained on most of the high-quality text that exists. Future models cannot simply be trained on "more" — there is no more, unless it is created or synthesised.

🔒
Crisis 2: Access and Privacy

The data that does exist — private enterprise data, medical records, legal documents, financial transactions — is behind walls of privacy regulation, competitive sensitivity, and contractual restriction. GDPR, HIPAA, CCPA, and equivalent regulations globally make accessing the most valuable training data legally complex and expensive. The richest data is the least accessible.

⚖️
Crisis 3: Copyright and Legal Risk

AI models trained on copyrighted content without licensing face existential legal liability. The Anthropic copyright settlement — $1.5 billion, 44% of the entire 2025 training data market — demonstrates the scale of the risk. News Corp sued, settled for $250 million over five years. Authors, record labels, image agencies, and news publishers are all in various stages of litigation or negotiation. The cost of legal data is now comparable to the cost of computing.

🧬
Crisis 4: Model Collapse

The most dangerous long-term risk: if AI models train on AI-generated content — their own outputs — errors and biases compound through successive generations, degrading model quality in a phenomenon called "model collapse." A 2024 Nature paper demonstrated this mathematically. As more of the internet fills with AI-generated text, the contamination of future training data becomes an existential quality problem. The internet is already estimated to contain 60%+ AI-generated content in some domains.

The Licensing Gold Rush — Who Is Buying What

If data is the oil,
the licensing deals
are the land grabs.
And the landowners
just discovered
what they're sitting on.

The recognition that high-quality human-generated text is scarce has created the fastest-moving licensing market in technology history. In 2024 alone, 91 major AI data licensing deals were disclosed. The pattern: AI companies buying permanent or time-limited rights to train on the content of publishers, news organisations, academic databases, and social media platforms that spent decades building the most valuable text corpus in existence — without knowing it was training data for AI.

Parties Value What Was Licensed · Why It Matters
OpenAI + News Corp
$250M / 5yr
Largest disclosed media licensing deal. Grants access to Wall Street Journal, Barron's, MarketWatch, New York Post, The Times (UK), The Sunday Times, The Sun, and dozens more. News Corp CEO: "an historic agreement that will set new standards for veracity, virtue, and value in the digital age." $50M+ annually — pure profit for News Corp at 2.5× their prior five-year net income.
OpenAI + Google + Reddit
$130M/yr
The most important real-time data deal. Reddit sells a structured real-time API stream rather than a historical archive — $70M from OpenAI, $60M from Google annually. Reddit's AI licensing revenue is now approximately 10% of total company revenue. Proof that social media platforms sitting on decades of human conversation are now a primary data asset class.
Anthropic settlement
$1.5B
44% of the entire 2025 training data market. Authors filed class action over unlicensed use of copyrighted books for training. The $1.5 billion settlement demonstrates that the legal liability for training on unlicensed content can approach the cost of AI computing infrastructure itself. Every AI lab has repriced their copyright risk after this settlement.
Wikimedia Enterprise
Ongoing
Amazon, Meta, Microsoft, Perplexity, Mistral all joined Wikimedia Enterprise in January 2026 — paying for "high-volume, high-speed access designed for AI" as a formal product, replacing free scraping. Bot traffic had driven multimedia bandwidth up 50% while human visits fell 8%. Wikimedia converted an AI-extracted resource into a paid product.
Scale AI + US DoD
$650M
The most significant government data deal. Scale AI secured a $650 million US Department of Defense contract in January 2026 for data labelling, annotation, and synthetic data generation for defence AI systems. Defence data — satellite imagery, signals intelligence, logistics — is the highest-value unlicensed data frontier. The government is now the largest single customer in the training data market.
Shutterstock AI licensing
$120M/yr
Shutterstock's AI training licensing revenue surpassed $120 million in 2025 — entirely from licensing its image archive to AI companies for training visual models. A stock photography company discovered its archive is worth more as AI training data than as stock photography. Getty Images, Adobe Stock, and other visual archives are following the same path.
The Solutions — Six Ways to Feed the AI

The wall is real.
The solutions are real too.
Each one creates
a new economic category.

Solution 1 Synthetic Data Generation Deployed · $0.92B Market 2026

If real data is scarce, manufacture it. Synthetic data — artificially generated datasets that mimic the statistical properties of real data — is the most powerful near-term solution to the data famine. Microsoft trained Phi-4, one of its most capable small language models, on 400 billion synthetic tokens with results comparable to models trained on much larger real datasets. Gartner predicted 75% of businesses would use synthetic data for AI training by 2026 — and the prediction looks conservative. The shift from 1% adoption in 2021 to 60% in 2024 is the fastest enterprise technology adoption in recent history.

The market is $0.92 billion in 2026, growing at 34.5% CAGR to $3 billion by 2030. Key players: NVIDIA (acquired Gretel AI for $320 million in March 2025, integrating into Omniverse Replicator with 252 enterprise deployments), MOSTLY AI (structured tabular data for financial services), MDClone (synthetic healthcare data), Synthesis AI, Tonic.ai. The domain where synthetic data works best: verifiable outputs — mathematics, code, formal reasoning — where models can generate solutions and check correctness automatically. The domain where it fails: creative writing, strategic planning, scientific hypothesis generation — wherever there is no objective verifier, errors compound into model collapse.

$0.92B → $3B by 2030. NVIDIA dominates after Gretel acquisition. Works brilliantly for verifiable domains. Dangerous for unverifiable ones.
Solution 2 Data Licensing Marketplaces Deployed · $4.8B Market 2025

The shift from free scraping to paid licensing is the most consequential economic restructuring in the content industry since the birth of digital advertising. The global dataset licensing for AI training market was $4.8 billion in 2025 and is projected to reach $22.6 billion by 2034. OpenAI alone has committed 53% of the disclosed deal value — approximately $816 million per year across 34+ deals.

The structural shift is from one-time training dumps to real-time API feeds. Reddit, Wikimedia, and Bloomberg are all moving toward continuous, paid, structured data access rather than historical archives. The most significant trend: content licensing is becoming real-time infrastructure, not historical asset sales. Scale AI established a dominant position with its $650 million DoD contract and its enterprise data labelling business. Appen, Lionbridge, and Sama provide the human annotation layer that makes raw data into training-quality data. The winner-take-most dynamic: $10 million is the minimum disclosed deal size, creating a market where only the largest publishers can participate in the first wave of value capture.

$4.8B in 2025, $22.6B by 2034. Scale AI + major publishers winning. Real-time feeds replacing one-time archives. Copyright liability repriced industry-wide.
Solution 3 Federated Learning & Privacy-Preserving AI Scaling · Unlocks Private Data

The richest data is behind privacy walls. Federated learning is the technique that unlocks it: instead of moving data to the model, the model moves to the data. A federated learning system trains on data where it lives — in a hospital, a bank, a retailer — without the raw data ever leaving the institution. Only model updates (gradients) are shared, not the underlying records. This makes GDPR, HIPAA, and CCPA compliance structurally embedded in the training process rather than a legal workaround.

Google has pioneered federated learning at scale — Android keyboard predictions are trained on device without user data leaving the phone. Apple uses federated learning across iOS devices. Healthcare is the highest-value application: hospitals with patient records containing 100 million medical imaging studies — among the most valuable training data for medical AI — can participate in model training without violating patient privacy. The economic implication: federated learning creates access to private enterprise data that would cost tens of billions of dollars to license conventionally.

Unlocks private healthcare, financial, and enterprise data. Google, Apple, Meta leading deployment. The solution that makes GDPR compliance an asset rather than a constraint.
Solution 4 Multimodal & Sensor Data Scaling · 10¹⁷ tokens/year available

When text runs out, use everything else. The International AI Safety Report 2026 quantified the data hierarchy: image data at 10¹⁴–10¹⁵ tokens, video at 10¹⁵–10¹⁶ tokens, IoT sensor data at 10¹⁷ tokens annually. The shift from text to multimodal training is the single biggest expansion of the addressable data universe available to AI developers — orders of magnitude larger than text. It requires new architectures (vision-language models, video understanding, sensor fusion) but the data itself is abundant.

Tesla's training approach demonstrates this at scale: the fleet of Tesla vehicles generates approximately 1.5 billion miles of real-world driving data annually — a multimodal sensor dataset (cameras, radar, ultrasound, GPS, accelerometers) that cannot be replicated through text or synthetic generation. That physical-world, embodied, sensor data is what trains autonomous driving beyond the capabilities of any text-based model. The companies that build the infrastructure to extract meaningful training signal from video, audio, and sensor data will have access to training data supplies measured in exabytes rather than terabytes.

Orders of magnitude more data than text. Requires new architectures to extract meaning. Tesla's driving data = the template. Video and IoT are the next frontiers.
Solution 5 Inference-Time Scaling & Chain-of-Thought Deployed · Changes the Equation

The most intellectually elegant solution to the data wall is not finding more data — it is getting more out of less. Inference-time scaling is the discovery that letting models "think longer" during use — generating and evaluating multiple reasoning paths before answering — produces dramatically better results without requiring more training data. OpenAI's o1 and o3 models, Anthropic's extended thinking, and Google's Gemini reasoning chains all demonstrate this: a model that reasons through a problem step-by-step outperforms a much larger model that doesn't, using the same training data.

The chain-of-thought breakthrough generates its own training data: models trained on millions of self-generated reasoning chains — where each step can be verified against known answers — create a virtuous cycle of self-improvement that does not require external data. This is the technique behind DeepSeek's efficient training, behind Phi-4's performance despite its small size, and behind the rapid capability improvements of 2025–2026 without proportional increases in training data consumption. Inference-time scaling effectively extends the useful life of existing training data by extracting more signal from it at deployment time.

Gets more from less. Chain-of-thought self-generates training data. o1, o3, extended thinking all demonstrate it. The most important architectural innovation of 2025.
Solution 6 Domain-Specific & Proprietary Data Moats Emerging · Most Durable Advantage

The long-term winner in the data economy will not be the company that finds the most data — it will be the company that builds the most valuable proprietary data moat. Bloomberg's financial data, Thomson Reuters' legal database, Epic Systems' electronic health records, Palantir's government intelligence data, Salesforce's CRM transaction data — these are datasets that cannot be scraped, purchased, or synthesised. They exist only because the company built the product that generated the data, and the data is only accessible to the company that owns it.

The economic logic: as public data becomes exhausted and licensed data becomes expensive, proprietary data becomes the primary source of competitive differentiation in AI. A model trained on Bloomberg's proprietary financial data will outperform a generic model on financial tasks not because it has more parameters but because it has better data. The implication for investors: companies with large, unique, defensible proprietary datasets are not primarily software companies or AI companies — they are data companies, and their data moats may be the most durable competitive advantages in the AI economy.

Bloomberg, Epic, Reuters, Palantir — data moats that cannot be bought or synthesised. The most durable AI advantage of the next decade. Proprietary data = competitive moat.
The Honest Read — Three Things That Make This Harder Than It Looks

Synthetic data solves the quantity problem but not the quality problem. The model collapse risk — AI systems trained on AI-generated content degrading in quality as errors compound — is mathematically real and empirically demonstrated. In domains with verifiable answers (maths, code), synthetic data works brilliantly. In domains without verifiable answers (creative writing, judgement, nuance), synthetic data degrades model quality. As more of the internet fills with AI-generated text, future models face the prospect of training on a corpus increasingly contaminated by the outputs of their predecessors. The 2024 Nature paper demonstrating model collapse is not a hypothetical warning — it is a mathematical proof with empirical confirmation. The data quality problem is harder than the data quantity problem, and solving quantity while ignoring quality is not a solution.

The winner-take-most dynamics in data licensing create a structural inequality that mirrors the content industry's historical pattern. Only the largest publishers — News Corp, Reuters, Shutterstock — have the market power to negotiate nine-figure licensing deals. The $10 million minimum disclosed deal size means that the majority of content creators — independent authors, individual journalists, small publishers, academic researchers — cannot participate in the licensing economy at all. Their work is used for training (or was, before the copyright litigation wave), but they receive no compensation. The Anthropic settlement's $1.5 billion is divided among a class of authors; individual shares will be small. The data licensing market is replicating the pattern of the music streaming revolution: the platforms capture the value, the creators receive fractions.

IoT sensor data is the largest untapped frontier — and also the most geopolitically contested. The 10¹⁷ tokens of annual IoT data includes factory sensors, smart city infrastructure, agricultural monitors, vehicle fleets, and energy grid sensors. Most of this data is owned by corporations or governments in countries with different data sovereignty rules. China's data laws prohibit export of certain categories of data; the EU's data governance framework restricts cross-border flows; India is building a data localisation regime. The geopolitics of data sovereignty are creating the same fragmentation in the data economy that sanctions and export controls are creating in the semiconductor economy — a world of data blocs rather than a global data commons.

The NGE View

The verdict.

What We Believe
The data famine is the most important constraint on AI progress that is not being adequately priced by public markets. AI companies are valued on their model capabilities and their revenue trajectories. They are not being valued on their data supply security — their ability to access the training data they will need for the next generation of models. As public text data approaches exhaustion and licensed data becomes expensive, the cost of training frontier models will increasingly be determined by data costs rather than compute costs. The companies that have secured proprietary data sources — through licensing deals, federated learning partnerships, or proprietary product data generation — have a competitive advantage that compounds over time and is invisible in current earnings multiples.
Synthetic data is the most immediately investable theme — and NVIDIA has already made the winning acquisition. NVIDIA's $320 million acquisition of Gretel AI in March 2025 — integrated into Omniverse Replicator with 252 enterprise deployments — positions it as the dominant infrastructure provider for synthetic data generation at the same moment the data famine makes synthetic data essential. The synthetic data market at $0.92 billion in 2026, growing at 34.5% CAGR to $3 billion by 2030, is being driven by the same urgency that drives AI compute spending. The picks-and-shovels analogy applies directly: when AI labs need synthetic data, they need NVIDIA's hardware to generate it and NVIDIA's tools to manage it.
Proprietary data moats are the most durable competitive advantage in the AI economy — and they are undervalued in companies that are perceived primarily as software or financial services businesses. Bloomberg Terminal's 320,000 users generate and consume the most valuable financial data in existence. Epic Systems' EHR platform covers 78% of US hospital patients. Palantir's government data platforms process intelligence that cannot be replicated. Thomson Reuters' Westlaw contains legal precedent going back centuries. These companies are not primarily software companies — they are data companies. And in a world where training data is scarce, their data moats are more valuable than their software businesses by a wide margin that the market has not yet fully recognised.
The shift from text to multimodal and sensor data is the next phase of the AI training revolution — and it favours physical-world companies over pure digital ones. When video and IoT sensor data become the primary training medium, the companies with the richest sensor data — Tesla (driving), Siemens (industrial), John Deere (agricultural), Medtronic (medical) — become more valuable as AI training data providers than as their primary businesses would suggest. The data famine does not end AI progress. It redirects it — from text-based intelligence to embodied, physical-world intelligence grounded in sensor data that only exists because companies built physical products in the real world. The data famine is not the end of AI. It is the gate through which only companies with proprietary physical-world data can pass.
NGE · A Futuristic Investment Letter

Long-horizon thinking on capital, technology, and the forces shaping the next decade of wealth creation. Written from first principles. Not consensus. Not noise.

— Pawan Bhatia · NextGen Economics · Bangalore, India