Tutorial
Source systems change constantly. A release adds a column, renames another, widens a type or splits a table. This is schema drift, and for test data management it is dangerous: policies were approved for the schema as it was, not as it is now.
Jobs also fail for ordinary reasons — a network drop, a full disk, a locked table. Recovery needs to be as deliberate as the happy path.
Text description
- Drift handling (Your environment): Detect drift → Compare to approved schema → Block uncertified output → Classify new fields → Resume or roll back
Why it matters
The most common way sensitive data leaks into test environments is not a dramatic breach; it is a new column nobody classified. If a table gains secondary_email and the masking policy only knows email, a naive tool copies the new column unmasked. Silent failures are just as harmful: a job that dies half-way and leaves a partial dataset in QA produces confusing defects and may bypass validation. Teams need drift to be caught before data moves, and failures to leave environments in a known, certified state.
Text description
- Gates (all must pass) (Your environment): Coverage → Masking verified → Referential integrity → Schema match → Quality rules → Provenance & checksums
Connection: Fail closed
- Outcome: All pass → certified, provisionable → Any fail → blocked, never provisioned
Example
A synthetic drift scenario on an invented customers table:
| Column | Approved schema | Current source | Action |
|---|---|---|---|
email | classified, masked | unchanged | proceed |
phone | classified, masked | type widened | re-validate rule |
secondary_email | — | new | block until classified |
fax | classified, masked | dropped | update policy |
The refresh is paused. Discovery classifies secondary_email as a direct identifier, a data owner approves policy version 8 that masks it, and the job resumes. Until then, QA keeps serving the last certified version.
How DataNivra approaches it
- Detect early. Each build starts with discovery; the agent compares the live schema with the one the policy was approved for.
- Fail closed. Unclassified or changed sensitive fields stop the job. Nothing uncertified is published.
- Classify, then approve. New fields go through Classification and policy approval before they are used.
- Recover cleanly. A failed job cleans its local workspace and reports an error code, never data. The previous certified version stays in place, and the next Refresh rebuilds from scratch.
- Record everything. Drift findings, pauses and approvals appear in the audit trail.
Only schema metadata — table and column names, types and counts — is reported to the control plane. See Refresh cadence and the interactive demo.
Key takeaways
- Treat every new column as sensitive until classified.
- Prefer a paused refresh over an unsafe one.
- Always fall back to the last certified version.