Learning center

Schema drift and failure recovery

How to handle source schema changes and failed jobs safely — detect drift, classify new fields first, block uncertified output and recover cleanly.

Tutorial

Source systems change constantly. A release adds a column, renames another, widens a type or splits a table. This is schema drift, and for test data management it is dangerous: policies were approved for the schema as it was, not as it is now.

Jobs also fail for ordinary reasons — a network drop, a full disk, a locked table. Recovery needs to be as deliberate as the happy path.

Schema drift and recoveryAt discovery time the agent detects drift and compares the schema with the approved one. Uncertified output is blocked, new fields are classified before use, and the job resumes or rolls back to the last certified version.Drift handling · Your environmentDetect driftCompare toapproved schemaBlockuncertifiedoutputClassify newfieldsResume or rollback
Schema drift and recovery. At discovery time the agent detects drift and compares the schema with the approved one. Uncertified output is blocked, new fields are classified before use, and the job resumes or rolls back to the last certified version.
Text description
  1. Drift handling (Your environment): Detect drift → Compare to approved schema → Block uncertified output → Classify new fields → Resume or roll back

Why it matters

The most common way sensitive data leaks into test environments is not a dramatic breach; it is a new column nobody classified. If a table gains secondary_email and the masking policy only knows email, a naive tool copies the new column unmasked. Silent failures are just as harmful: a job that dies half-way and leaves a partial dataset in QA produces confusing defects and may bypass validation. Teams need drift to be caught before data moves, and failures to leave environments in a known, certified state.

Dataset certification gatesA dataset must pass six gates: coverage, masking verification, referential integrity, schema match, quality rules, and provenance with checksums. If every gate passes the dataset is certified and may be provisioned. If any gate fails, the dataset is blocked and is never provisioned.Gates (all must pass) · Your environmentCoverageMaskingverifiedReferentialintegritySchema matchQuality rulesProvenance &checksumsOutcomeAll pass → certified,provisionableAny fail → blocked,never provisionedFail closed
Dataset certification gates. A dataset must pass six gates: coverage, masking verification, referential integrity, schema match, quality rules, and provenance with checksums. If every gate passes the dataset is certified and may be provisioned. If any gate fails, the dataset is blocked and is never provisioned.
Text description
  1. Gates (all must pass) (Your environment): Coverage → Masking verified → Referential integrity → Schema match → Quality rules → Provenance & checksums

    Connection: Fail closed

  2. Outcome: All pass → certified, provisionable → Any fail → blocked, never provisioned

Example

A synthetic drift scenario on an invented customers table:

ColumnApproved schemaCurrent sourceAction
emailclassified, maskedunchangedproceed
phoneclassified, maskedtype widenedre-validate rule
secondary_email—newblock until classified
faxclassified, maskeddroppedupdate policy

The refresh is paused. Discovery classifies secondary_email as a direct identifier, a data owner approves policy version 8 that masks it, and the job resumes. Until then, QA keeps serving the last certified version.

How DataNivra approaches it

  • Detect early. Each build starts with discovery; the agent compares the live schema with the one the policy was approved for.
  • Fail closed. Unclassified or changed sensitive fields stop the job. Nothing uncertified is published.
  • Classify, then approve. New fields go through Classification and policy approval before they are used.
  • Recover cleanly. A failed job cleans its local workspace and reports an error code, never data. The previous certified version stays in place, and the next Refresh rebuilds from scratch.
  • Record everything. Drift findings, pauses and approvals appear in the audit trail.

Only schema metadata — table and column names, types and counts — is reported to the control plane. See Refresh cadence and the interactive demo.

Key takeaways

  • Treat every new column as sensitive until classified.
  • Prefer a paused refresh over an unsafe one.
  • Always fall back to the last certified version.