Tutorial
There are two broad ways to get realistic data into a test environment. You can start from production, take a subset and mask it, or you can generate records from rules and models without reading production at all. Both are legitimate, and each has blind spots.
The choice is rarely either-or. The useful question is which tests need which kind of data, and what privacy posture each environment requires.
Text description
- Masked production-like (Your environment): Real distributions → Real edge cases → Needs source access (in your environment) → Privacy via masking policy
- Fully synthetic: Generated from rules/models → No source rows needed → Scenarios on demand → Fidelity depends on generator
Why it matters
Masked production-like data carries the messiness of real life: odd distributions, legacy formats, historical records that nobody would think to invent. That is exactly what finds certain defects. But it requires reading production, which must happen inside your environment under policy, and masking must be done carefully to keep privacy risk low.
Synthetic data needs no production rows at all, can create scenarios that production does not contain yet, and is easy to share widely. Its weakness is fidelity: it only contains the patterns its generator knows about. Choosing badly leads either to unnecessary exposure or to tests that pass on data that never looks like reality.
Example
A synthetic health-plan team (every name and number here is invented) needs data for three purposes:
| Need | Better fit | Reason |
|---|---|---|
| Regression tests on claim adjudication | Masked subset | Real code combinations and edge cases matter |
| A new feature for a plan type that launches next quarter | Synthetic | No production records exist yet |
| A public demo environment | Synthetic | Nothing derived from real members should appear |
For the regression suite, the team subsets a few thousand members with their claims and masks identifiers deterministically. For the new plan type, it generates members from pack templates with scenarios such as a mid-year plan change. The demo uses synthetic data only.
How DataNivra approaches it
DataNivra supports both paths in the same pipeline, and both run as customer-resident processing. For masked data, the agent reads sources read-only, builds a subset and applies an approved masking policy; for synthetic data, the engine generates records from industry-pack templates and scenarios. The two can be mixed, for example a masked subset enriched with synthetic edge cases.
Whichever path you pick, the result goes through the same certification gates before it can be provisioned, and synthetic records are labelled as such so they are never confused with real ones. Masking is not treated as proof of Anonymization; policies are designed and reviewed as pseudonymization unless stronger techniques are applied. Only metadata and aggregate metrics reach the control plane, preserving zero raw-production-data egress. Try both paths in the synthetic interactive demo.
A quick decision guide
- Testing a new feature with no production data yet? Synthetic. There is nothing to mask.
- Reproducing a production defect? Masked production-like data, subset around the affected entities, so the same shape of data is present.
- Performance and volume testing? Usually synthetic at the target volume, or a masked subset scaled up with synthetic rows.
- Negative and boundary testing? Synthetic, with rows explicitly labelled as negative-test data so broken records never masquerade as normal ones.
- Developer laptops and shared demos? Synthetic by default; it reduces what a lost laptop can expose.
- UAT with business users? Masked data often works best, because users recognise realistic patterns.
Many teams combine both: a masked subset for realism plus synthetic rows for the scenarios production has never seen. The synthetic-scenario planner among the free tools helps you choose the mix, and masking vs tokenization compares masking techniques. See which sources can feed masked data on the integrations page, read the documentation, or start free with the synthetic sandbox.
Key takeaways
- Masked data brings realism; synthetic data brings freedom and new scenarios.
- Match the data type to the test and the environment.
- Certify both the same way before use.