Banks Bet on Synthetic Data to Train AI — But Is It Safe?
From clean rooms to simulated customers, financial firms are racing to create usable datasets for generative AI while dodging privacy pitfalls
From clean rooms to simulated customers, financial firms are racing to create usable datasets for generative AI while dodging privacy pitfalls

Illustration by IMF Alpha editorial · Reviewed by Pedro Marini
Banks and fintechs are quietly swapping real ledgers for simulated ones
As teams scramble to tune generative models, a new supply-chain headache has emerged. Real customer data is toxic to share, yet models still need believable examples. Enter synthetic data: algorithmically produced records that mimic customer behavior without carrying real identities. For U.S. financial firms the appeal is obvious — faster model development, fewer compliance battles, and a way around long, expensive anonymization efforts. But the trade-offs matter.
Why now
Momentum is real. Maturity is another question. Think of synthetic datasets as stunt doubles: convincing from a distance, but they must not be mistaken for the real person in a close-up. That distinction is precisely where the risk lives.
Practical trade-offs
A short history lesson
Anonymization has failed before. The internet learned that the hard way when supposedly scrubbed search logs were deanonymized. Financial institutions face higher stakes — reputational damage, regulatory fines, and consumer suits — so the margin for error is thin.
Where firms are splitting
Watch for
Regulators will eventually spell out rules around synthetic data; compliance teams should not treat it as a free pass. Independent metrics will matter — statistical distance is only the beginning. Real safety comes from stress tests that probe edge cases and adversarial attempts to reidentify data.
A practical roadmap for executives
Synthetic data is not a cure-all. It is, however, a useful tool — if banks build the guardrails so those simulated customers never become real headaches. Expect surprises; plan for them.
Pedro Marini

Smartphones and PCs are starting to run generative models locally. That shifts power to chipmakers, changes app economics, and gives privacy a new marketing lifeline.

From privacy-by-default budgeting to instant fraud checks, on-device generative models are reshaping fintech. Here’s what consumers, banks and investors should watch next.

A new wave of phone fraud uses synthetic voices to bypass agents and customers. Financial firms pivot to biometrics, behavioral signals and stricter verification.