The New Gold Rush: Paying for Data to Power AI
After years of free-for-all scraping, companies are buying and synthesizing datasets. That shift is creating winners — and a new investment playbook.
After years of free-for-all scraping, companies are buying and synthesizing datasets. That shift is creating winners — and a new investment playbook.

Illustration by IMF Alpha editorial · Reviewed by Pedro Marini
The shift is clear: data is no longer something you scrape and forget.
Over the past 18 months the economics and legal realities around training large language models have pushed firms toward curated, licensed and synthetic datasets. It may sound like an operational tweak, but it changes a central input for AI: data that is reliable and auditable.
Why this matters now
A short history
Ten years ago most teams hoarded web crawls and hoped for the best. That gave scale and speed but left big blind spots: legal exposure, skewed samples, labels you could not trust. The last few years have felt like a correction — the industry moving from wildcatting toward more structured, accountable data practices.
Who’s gaining ground (and why it matters)
Investors: a few practical angles
Risks and caveats
A final take
We are shifting from a web-scrape era to a paid-and-proven data era. That favors companies that can provide governed access, audit trails and privacy-preserving alternatives. For investors the sensible bet is on platforms that combine distribution with trust services — and to be skeptical of expensive datasets sold without clear ROI.
My read: this is less a temporary fad and more a structural change in where value accrues in the AI stack. Treat data as the product, not a free input.

OpenAI's enterprise revenue has reportedly surpassed $2 billion annually, signaling rapid adoption of its AI services by businesses and solidifying its market position.

Recent fintech earnings reports emphasize the critical role of payment processing volumes and the emerging impact of AI-driven underwriting models on profitability.

Asset managers and hedge funds are quietly building proprietary data lakes to train in-house AI — reshaping competitive moats, privacy risks, and market structure.