Learning center

Data masking

How masking replaces sensitive values with realistic substitutes, why determinism matters, and how to keep masked data useful for testing.

Tutorial

Masking transforms sensitive values into substitutes that look and behave like real data but do not reveal the originals. A good masked dataset still exercises the same code paths — validation, joins, sorting, formatting — as production data would.

Masking is a family of techniques rather than one: substitution from realistic dictionaries, shuffling within a column, date shifting, format-preserving replacement, nulling or suppression, and generalisation such as reducing a birth date to a year.

Masking a synthetic recordA synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.Before (synthetic example)name: Jordan Blakemember_id: M-40021dob: 1984-03-17email:jordan.b@example.testAfter masking · Your environmentname: Riley Chenmember_id: M-91877dob: 1984-05-02email:user_7f3a@example.testApproved masking policy, keyed & deterministic
Masking a synthetic record. A synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.
Text description
  1. Before (synthetic example): name: Jordan Blake → member_id: M-40021 → dob: 1984-03-17 → email: jordan.b@example.test

    Connection: Approved masking policy, keyed & deterministic

  2. After masking (Your environment): name: Riley Chen → member_id: M-91877 → dob: 1984-05-02 → email: user_7f3a@example.test

Why it matters

Naive masking breaks tests. Replace every name with XXXX and name-matching logic is never exercised. Randomise member identifiers differently in each table and joins return nothing. Make card numbers fail their check digit and payment validation rejects every record.

On the privacy side, weak masking gives false comfort. Leaving quasi-identifiers untouched, or masking with an unkeyed hash that anyone can recompute, may let someone reverse the result. Masking needs to be both useful for testing and deliberate about what it protects.

Example

A synthetic member record passes through a masking policy:

FieldBefore (synthetic)AfterTechnique
nameJordan BlakeRiley ChenDictionary substitution
member_idM-40021M-91877Keyed deterministic, same format
dob1984-03-171984-05-02Per-member date shift
emailjordan.b@example.testuser_7f3a@example.testNon-deliverable substitute

Because masking of member_id is deterministic, M-40021 becomes M-91877 in the claims table, the eligibility table and the downstream data warehouse alike. The shift applied to the birth date is applied consistently to that member's other dates so intervals between events are preserved.

How DataNivra approaches it

Masking policies in DataNivra are versioned, reviewed and approved in the control plane, and shipped to the agent as a snapshot with a checksum that the agent re-verifies before running. The masking itself is customer-resident: the agent applies it inside your environment, and masking keys stay in your own secret store, referenced by name only.

Rules support deterministic keyed transformations, format-preserving substitution, date shifting and free-text suppression, with presets from industry packs that you adopt as your own policy. After masking, certification verifies that the policy was actually applied to every classified column; a dataset that fails is never provisioned. See Dataset certification.

Key takeaways

  • Masked data must still behave like real data.
  • Determinism keeps joins and cross-system links intact.
  • Keep keys out of the masked data's reach, and verify masking before use.