Tutorial
Most organisations run a chain of non-production environment tiers. Each exists for a different kind of confidence, and each needs different data. Treating them all the same — usually by giving every environment the same copy — wastes money and multiplies risk.
This tutorial walks up the ladder from development to performance testing and describes what "good data" means at each rung.
Text description
- Environments: DEV: small & fast → QA: representative subset → SIT: consistent across systems → UAT: realistic scenarios → PERF: scale & distribution
Why it matters
When every environment gets a full copy, developers wait for huge restores they do not need, integration teams fight over shared records, and performance testers still lack the volume they do need. Matching data to purpose shortens feedback loops and shrinks the footprint of sensitive data.
It also clarifies ownership. A UAT dataset built around business scenarios has different approvers than a performance dataset sized for load. Naming those differences makes requests easier to automate and approve.
Example
A synthetic retail bank maps its environments like this:
| Tier | Purpose | Data shape |
|---|---|---|
| DEV | Fast feedback for engineers | Small, often fully synthetic (Synthetic data), rebuilt on demand |
| QA | Functional regression | Representative masked subset (Subsetting), stable between runs |
| SIT | Systems working together | Same entities across core banking, cards and CRM, with Referential integrity across systems |
| UAT | Business sign-off | Recognisable business scenarios: overdrafts, disputes, joint accounts |
| PERF | Load and capacity | Production-like volume and distribution, masked or synthesized at scale |
The bank, systems and scenarios are illustrative.
Note how SIT is the hardest: a customer masked one way in the card system and another way in CRM breaks every integration test.
How DataNivra approaches it
In DataNivra a dataset request names its target environment, and policy can differ per tier — for example, DEV may allow only synthetic data, while QA allows masked subsets with approval. The agent builds each dataset inside your environment and provisions it to the named target.
Deterministic, keyed masking gives the same substitute for the same input in every system, which is what makes SIT data line up. Performance datasets can combine masked subsets with synthetic volume. Every dataset is certified before provisioning and carries its own refresh and retention settings. See How it works or try the interactive demo.
Key takeaways
- Different environments answer different questions; give them different data.
- Cross-system consistency matters most in SIT.
- Encode per-environment rules in policy, not in tribal knowledge.