Learning center

Dataset certification

What it means to certify a test dataset, which gates it must pass, and why a failed gate must block provisioning rather than warn.

Tutorial

Masking a dataset is an action. Certifying it is a claim: this dataset meets the policy it was built under, and here is the evidence. The distinction matters because many things can go wrong between "we ran the masking job" and "the data is safe and usable".

Certification turns that claim into a set of automated checks, each with a pass or fail outcome and a record of what was checked.

Dataset certification gatesA dataset must pass six gates: coverage, masking verification, referential integrity, schema match, quality rules, and provenance with checksums. If every gate passes the dataset is certified and may be provisioned. If any gate fails, the dataset is blocked and is never provisioned.Gates (all must pass) · Your environmentCoverageMaskingverifiedReferentialintegritySchema matchQuality rulesProvenance &checksumsOutcomeAll pass → certified,provisionableAny fail → blocked,never provisionedFail closed
Dataset certification gates. A dataset must pass six gates: coverage, masking verification, referential integrity, schema match, quality rules, and provenance with checksums. If every gate passes the dataset is certified and may be provisioned. If any gate fails, the dataset is blocked and is never provisioned.
Text description
  1. Gates (all must pass) (Your environment): Coverage → Masking verified → Referential integrity → Schema match → Quality rules → Provenance & checksums

    Connection: Fail closed

  2. Outcome: All pass → certified, provisionable → Any fail → blocked, never provisioned

Why it matters

Without certification, teams rely on trust. A job finished, so the data is assumed to be fine. But a new column may have arrived unmasked, a filter may have dropped parent records, or a transformation may have produced values the application rejects. Each of these reaches testers as a real problem: exposure of sensitive values, broken tests, or wasted time. Certification catches them before provisioning. It also gives security and data owners something concrete to review, instead of a verbal assurance, and it gives auditors a trail that links a dataset version to the policy and checks behind it.

The DataNivra pipeline, end to endAll eight stages run inside your environment. Preparation: discover, classify, subset, mask. Proof and delivery: synthesize, validate, certify, provision. Only status, counts and evidence references are reported to the control plane.Prepare · Your environmentDiscoverClassifySubsetMaskProve & deliver · Your environmentSynthesizeValidateCertifyProvisionSame job, same environment
The DataNivra pipeline, end to end. All eight stages run inside your environment. Preparation: discover, classify, subset, mask. Proof and delivery: synthesize, validate, certify, provision. Only status, counts and evidence references are reported to the control plane.
Text description
  1. Prepare (Your environment): Discover → Classify → Subset → Mask

    Connection: Same job, same environment

  2. Prove & deliver (Your environment): Synthesize → Validate → Certify → Provision

Example

A synthetic QA dataset for an invented retailer is built from 800 customers. The certification run reports:

GateResultDetail (aggregate only)
Coveragepassall requested entities present
Masking verifiedfail1 column classified sensitive had no masking rule
Referential integritypass0 dangling keys
Schema matchpassmatches approved schema version
Quality rulespass0 rule violations
Provenance and checksumspassmanifest checksum recorded

One gate failed, so the dataset is not certified. It is not provisioned to QA, and the owner is told which column and which rule are missing, never the values in it. After the policy is updated and approved, the job reruns and all gates pass.

How DataNivra approaches it

Certification runs inside your environment as the last step before provisioning. The gates cover coverage, masking, referential integrity, schema, quality and provenance. The behaviour is Fail closed: if any gate fails, or if the outcome is uncertain, the dataset is marked failed and never provisions. A revoked dataset cannot be provisioned either.

Evidence such as the certification report and dataset manifest stays in your environment. The control plane stores an Evidence reference — a URI and checksum — plus gate outcomes and aggregate metrics, which keeps zero raw-production-data egress intact while giving the console a complete status view. Provisioning only ever sees certified versions. See how it works for where certification fits in the flow.

Reading a certification result

  • Every gate is pass or fail. There is no partial pass. If one gate fails, the dataset is not certified and cannot be provisioned.
  • Policy coverage and masking completion show whether every sensitive column the policy covers was actually transformed.
  • Referential integrity and orphan detection show whether every reference still resolves after subsetting and masking.
  • Row-count reconciliation and schema validation catch truncated loads and schema drift.
  • Provenance, versions and checksums record which policy and engine produced the data and let anyone confirm later that the files were not changed.

When a gate fails, fix the cause rather than the symptom: add the missing column to the policy, declare the missing relationship, or refresh after a schema change. A refusal is the system working as intended, not an obstacle to route around. Certification is also not permanent: if a dataset is later revoked, for example because a policy was found to be incomplete, it can no longer be provisioned and environments should move to a newly certified version.

The guides explain certification for specific sources, the integrations page lists supported connectors, and the documentation details the evidence manifest. Plan certification into your pipeline with the checklists among the free tools, or start free and certify a synthetic dataset yourself.

Verify it yourself: make certification fail on purpose

A certification claim is only worth something if you can watch it refuse. The interactive demo runs on a synthetic estate in your browser and lets you break a dataset deliberately:

  1. Run discovery and classification, choose a subset, and at the masking step tick Leave a column unmasked for one sensitive column, for example the e-mail address.
  2. Apply masking and run certification. The masking gate fails, its result names the unmasked column — never the values in it — and the dataset is marked failed.
  3. Try to provision it. Provisioning is refused: a failed dataset can never reach a test environment, whatever button is pressed.
  4. Start the demo again, leave every column masked and certify. Every gate passes and provisioning is allowed.

In your own environment the same check is worth running once as an acceptance test: remove a sensitive column from an approved policy on a synthetic source, run the job, and confirm that the dataset fails certification and that nothing is provisioned. The same rule covers the whole requested dataset: when a request spans several systems, every table from every source goes through the same gates, and one failing table fails the dataset.

Key takeaways

  • Masked is not the same as certified.
  • Every gate must pass; a failure blocks provisioning.
  • Evidence stays with you; the cloud keeps references.