S&P 5005,842.10 0.42%
NASDAQ19,210.55 0.88%
NVDA1,184.22 2.41%
MSFT478.90 0.88%
GOOGL210.11 1.12%
META612.50 0.34%
AAPL239.80 0.21%
AMZN248.66 1.40%
AVGO1,902.40 3.12%
TSLA298.10 1.05%
BTC98,420 1.88%
ETH4,210 2.24%
10Y4.18% 0.02%
DXY104.12 0.18%
S&P 5005,842.10 0.42%
NASDAQ19,210.55 0.88%
NVDA1,184.22 2.41%
MSFT478.90 0.88%
GOOGL210.11 1.12%
META612.50 0.34%
AAPL239.80 0.21%
AMZN248.66 1.40%
AVGO1,902.40 3.12%
TSLA298.10 1.05%
BTC98,420 1.88%
ETH4,210 2.24%
10Y4.18% 0.02%
DXY104.12 0.18%
Back to homepage
Synthetic Data

Synthetic Data Is Eating the Training Set: How Wall Street Should Reprice AI Bets

From data marketplaces to GPU demand, a quiet supply shock in training data is shifting winners in the AI race — and not always in predictable ways.

P
Pedro Marini
August 6, 2026 · 4 min read
Synthetic Data Is Eating the Training Set: How Wall Street Should Reprice AI Bets

Illustration by IMF Alpha editorial · Reviewed by Pedro Marini

Listen to this article
AI narration · ~4 min
Tickers mentioned
NVDA+2.30%SNOW+1.80%PLTR+3.50%MSFT-0.80%GOOG+1.10%AMZN+0.50%META-1.20%

The fast money in AI so far has gone to model scale and raw compute. But there’s a quieter supply shock taking shape that could matter more over time: synthetic data is becoming the preferred feedstock for training and fine-tuning large models. That subtle shift rearranges the value chain — think cloud marketplaces and data integrators, not just GPUs — and it changes how investors should price risk and opportunity.

Why synthetic data is starting to scale

  • Privacy and regulation are tightening. Companies can’t keep hoarding user logs the way they used to; synthetic alternatives offer a cleaner legal posture and less friction.
  • Labeling is expensive and slow. Human-in-the-loop costs, plus the rarity of edge cases, make some datasets impractical to assemble. Synthetic data can simulate those hard-to-observe events at far lower marginal cost.
  • Speed beats tiny gains in accuracy. Teams can spin up dozens or hundreds of synthetic variants and iterate quickly — which matters more for product velocity than squeezing out the last fraction of performance.

What’s interesting is how this looks in practice: synthetic data doesn’t replace real-world signals, it amplifies experimentation. A useful analogy is fertilizer — it boosts yield and shortens seasons. But misapplied, it can create brittle blooms that don’t survive when conditions change.

Who stands to gain — and why some names are underrated

  • Cloud and data marketplaces become natural aggregators. Snowflake, for example, is already positioning its marketplace as a channel for synthetic datasets and labeling services. That could be a steady, high-margin revenue stream if adoption follows.
  • GPU demand won’t vanish, but it will morph. Expect fewer gargantuan pre-training runs and more frequent, intense fine-tuning bursts. That pattern helps both major GPU vendors and companies focused on efficient inference and tuning pipelines.
  • Data integration, provenance, and observability vendors win if they can prove quality. In regulated industries especially, firms that can tie synthetic datasets to tests, drift monitoring, and clear provenance become hard to replace.

Notice the asymmetry: flashy model vendors get attention, but the quiet middleware — the systems that make synthetic data trustworthy and operational — may capture the most durable value.

The counterarguments investors should weigh

  • Distributional mismatch. Synthetic samples can miss or distort real-world correlations. Models trained on them can be brittle when reality diverges from the generator.
  • IP and provenance issues. If synthetic data is generated from models trained on proprietary corpora, disputes will shift from outputs back to training assets. Legal gray areas are inevitable.
  • Concentration risk. If a few firms control most synthetic pipelines, customers face pricing power and single-point failures. That creates systemic vendor risk for enterprises.

In short: promising, yes — but messy in practice. Some teams underestimate how often synthetic data introduces subtle biases that only show up in production.

Signals worth watching

  • Changes in revenue mix at Snowflake and other cloud leaders: accelerating marketplace or data-product monetization is a clear sign.
  • GPU order patterns shifting away from massive pre-training clusters toward more frequent, smaller fine-tune jobs. Read capex commentary from cloud providers and chip makers closely.
  • Partnerships between synthetic-data startups and regulated incumbents in healthcare, fintech, and automotive. Those ties usually mean stickier, higher-value adoption.

A concise investment thesis

Synthetic data will not make real-world data irrelevant. But it will compress iteration time and lower the marginal cost of experimentation. For equity investors that argues for looking beyond headline model vendors to the companies that control distribution, provenance, and integration — the plumbing that makes synthetic datasets reliable and actionable.

My read: the biggest winners are likely to be the unglamorous middleware plays and marketplaces that scale synthetic data responsibly, not necessarily the flashiest model builders. Call it unromantic — markets, more often than not, pay for the pipes, not the poetry.

Advertisement
Continue reading

Related coverage

The IMF Brief · Daily Newsletter

The AI economy, decoded before the open.

Five minutes. One email. The signal cutting through the noise at the intersection of artificial intelligence and Wall Street. Free, forever.

Join 184,000+ readers · No spam · Unsubscribe anytime