Synthetic vs masked production-like data
When to mask a governed subset of production, when to generate synthetic records, and why many teams combine both.
Topics
Synthetic data contains no real person at all, which makes it easy to share; masked data keeps the quirks of real records, which makes it better at finding real defects. Most teams need both. This hub compares the two and helps plan the synthetic scenarios — edge cases, negative tests, volume — that production data rarely provides.
A common pattern is a hybrid dataset: a masked, relationship-safe subset of real records for realism, topped up with generated records for the cases production never contains, such as an expired card, a claim filed before its policy began or a customer with ten thousand orders. Industry packs ship scenario templates of this kind for regulated domains.
When to mask a governed subset of production, when to generate synthetic records, and why many teams combine both.
The practical difference between pseudonymized and anonymized data, why the distinction matters for test data, and how to choose.
When a test should run on synthetic data and when it needs a masked subset of real data, with a runnable check of the markers that make synthetic rows unmistakable.
Inventory the denials, pended claims, out-of-range labs, encounter types and refills present in a synthetic healthcare bundle, and map them to the pack's named scenarios.
Count the payment failures, chargebacks, fraud cases and account states in a synthetic banking bundle and check that every fraud case points at a real transaction.