Use case
"Should we use synthetic data or masked production data?" is usually asked as if one answer had to win. In practice a mature test estate uses both, and the useful question is narrower: for this particular test, which source gives realistic behaviour with the least exposure? This page gives a rule of thumb you can argue with and a runnable check that shows how DataNivra's synthetic data is built so that nobody ever mistakes it for the real thing.
The scenario
A team is planning test data for a claims platform. Some tests are demos and public bug reproductions; some exercise rare paths such as denied claims and boundary lab values; some are performance runs; and a few regressions only reproduce with the quirks of real records — odd encodings, skewed distributions, historical data written by a system retired years ago. One source cannot serve all of them well.
Synthetic data is generated from a model of the schema and its business rules. Nothing in it came from a person, so it can go anywhere: laptops, public CI, a vendor's support ticket. It can contain rare cases on demand and in any volume. Its weakness is fidelity: it only knows the patterns somebody modelled.
A masked subset starts from real rows inside your network, keeps a representative slice with its relationships intact, and transforms every sensitive value. It carries the real distributions and the real mess. Its weakness is that it is still derived from people's records, so it needs policies, controls and evidence.
Run it
The check below reads the synthetic member file from the healthcare pack and counts the markers that make each row recognisably synthetic, then prints the rule of thumb.
import csv
import io
import os
import re
import urllib.request
BASE = os.environ.get("DATANIVRA_DOWNLOADS", "https://www.datanivra.com/downloads")
BUNDLE = f"{BASE}/packs/healthcare/2.1.0/synthetic"
with urllib.request.urlopen(f"{BUNDLE}/enrollment.members.csv") as response:
members = list(csv.DictReader(io.TextIOWrapper(response, encoding="utf-8")))
# Synthetic values are built to be unmistakable: reserved domains, fictional number ranges,
# a "ZZ" record-number prefix and a provenance column on every row.
markers = {
"email on a reserved example domain": lambda m: re.search(r"@example\.(com|org|net)quot;, m["email"]),
"phone in the fictional 555 range": lambda m: m["phone"].startswith("555-"),
"record number with the ZZ prefix": lambda m: m["medical_record_number"].startswith("MRN-ZZ-"),
"row marked SYNTHETIC": lambda m: m["_dn_provenance"] == "SYNTHETIC",
}
for label, test in markers.items():
present = [m for m in members if m["email"]] if label.startswith("email") else members
print(f"{label}: {sum(1 for m in present if test(m))} of {len(present)}")
# Which data a test needs decides the source: a rough rule of thumb, not a policy engine.
needs = [
("demo, sandbox or public bug report", "synthetic"),
("rare case missing from production (denials, boundary values)", "synthetic"),
("load test larger than any extract", "synthetic"),
("regression on real value distributions", "masked subset"),
("defect only seen with production records", "masked subset"),
("migration rehearsal on the real schema and volumes", "masked subset"),
]
for need, source in needs:
print(f"{source:>13} <- {need}")Expected output
email on a reserved example domain: 4 of 4
phone in the fictional 555 range: 5 of 5
record number with the ZZ prefix: 5 of 5
row marked SYNTHETIC: 5 of 5
synthetic <- demo, sandbox or public bug report
synthetic <- rare case missing from production (denials, boundary values)
synthetic <- load test larger than any extract
masked subset <- regression on real value distributions
masked subset <- defect only seen with production records
masked subset <- migration rehearsal on the real schema and volumesOne member in the sample has no email address at all — optional columns are left empty in realistic proportions, which is itself something tests should cover.
Schema
| Column | Synthetic convention |
|---|---|
email | addresses on reserved example.com/.org/.net domains, or empty |
phone | the fictional 555- range |
medical_record_number | MRN-ZZ- prefix, never a real issuer format |
_dn_provenance | SYNTHETIC on every generated row |
The bundle's SYNTHETIC_MANIFEST.json records the generator seed, scale and per-file row counts, so the same files can be regenerated and compared.
What DataNivra does with your own data
Both paths run in the same product. Synthetic generation fills scenarios from the industry pack's models and can top up a masked subset with rare cases it lacks, tagging generated rows in provenance. Masked subsets are built by the agent inside your network from a read-only source connection. Either way the result goes through the same validation and certification gates before it can be provisioned, and the hosted control plane never receives row values. Certification here means DataNivra-certified against configured policy gates; it is not a regulatory or third-party certification and does not make a system compliant with any regulation.
Limits to plan around
- Synthetic data reflects the model, not your production quirks. Performance results on it say little about real data skew.
- A masked subset of real data is still derived from personal records; treat its environments accordingly and keep its policies reviewed.
- The rule of thumb above is a starting point, not a classification of your tests, and the marker check covers one table of one pack.
Next steps
Read synthetic vs masked data, plan rare cases with the synthetic scenario planner, and try both flows in the synthetic playground.
Synthetic downloads
Files from the Healthcare & Life Sciences pack 2.1.0 bundle. Everything in it is synthetic, generated from a fixed seed, and listed with its SHA-256 digest in the bundle's MANIFEST.json.
Industry pack: Healthcare & Life Sciences — its entities, scenarios, policy templates and the complete asset bundle.
Try it with DataNivra
The synthetic playground walks through discovery, subsetting, masking, validation and certification in your browser, with no account. Starting free gives you the hosted synthetic sandbox; your own sources need an agent in your network.