Guides

Masking vs tokenization for test data

The practical difference between data masking and tokenization, when each fits test data, and how DataNivra implements both without sending values out of your network.

Guide

"Masking" and "tokenization" are often used as if they meant the same thing. They overlap, but they solve different problems, and choosing the wrong one for test data leads either to useless datasets or to datasets that are less protected than they look. This guide explains the difference in practical terms and shows how DataNivra implements each.

Definitions that actually help

Masking replaces a sensitive value with a different value that is fit for purpose. The replacement might be a realistic synthetic name, a shifted date, a perturbed amount, a redaction marker or a keyed pseudonym. The defining property is that the original is not recoverable from the masked dataset, and the replacement is chosen to keep the data useful: formats, lengths, distributions and joins.

Tokenization replaces a value with a token, an opaque surrogate such as tok_..., and typically keeps a mapping so that authorised systems can relate tokens back to the original. Tokenization was designed for production: a payment system stores tokens instead of card numbers, and only a hardened vault can map them back. The defining property is the vault, and the vault is the most sensitive system in the design.

What each means for test data

Test environments need data that behaves like production. That usually means:

  • values that pass validation (an email that looks like an email, a postal code in the right format);
  • identifiers that are unique and join consistently across tables and systems;
  • realistic spreads of dates and amounts, so business rules fire.

Masking is built for this. Tokens, by contrast, are opaque: a tokenized email fails validation, a tokenized date is not a date, and tests that parse or display these fields break. Tokenization is useful in test data for fields whose only job is to be a stable, unique reference, such as an external account reference that tests only compare for equality.

The bigger concern is reversibility. A reversible token vault reachable from test environments turns every test dataset back into production data for anyone who can call the vault. For test data, prefer irreversible transformations, and if you must keep a mapping, keep it inside the production boundary and never in lower environments.

How DataNivra implements masking

Masking policies assign a strategy to each classified column. Strategies include synthetic replacements for names, emails, phone numbers and addresses; date shifting that uses the same offset for the same entity so intervals stay realistic; numeric perturbation; redaction and nulling; hashing; keyed HMAC pseudonymisation; and a format-preserving permutation for bounded integer domains that is collision-free, so a masked primary key remains unique. Deterministic strategies use a key held in your secret store, which means the same customer id becomes the same pseudonym in every table, every system and every refresh, and referential integrity survives masking.

All masking runs in the agent inside your network. The key is referenced by name, never sent to DataNivra Cloud. Certification then checks that every column the policy requires to be masked actually was, and blocks the dataset otherwise.

How DataNivra implements tokenization

The TOKENIZE strategy uses a pluggable token vault. Two ship with the engine:

  • a stateless HMAC vault that produces tok_ plus a keyed digest: deterministic, consistent across jobs that share the key, irreversible, with nothing stored;
  • a customer-local SQLite vault that issues random tokens and persists the mapping so tokens stay stable across runs. The mapping is keyed by a keyed fingerprint of the value, not the value itself, so the vault file holds no plaintext source values.

Customers can register an adapter to an existing enterprise tokenization service through the plugin registry. Any such vault stays inside your boundary.

Choosing between them

FieldGood choiceWhy
Customer name, email, phoneSynthetic replacementTests need valid, realistic formats
Customer or account id used in joinsDeterministic pseudonym or format-preserving permutationUnique, consistent across tables
External reference compared only for equalityTokenOpaque is fine; stable across runs
Date of birth, admission dateEntity-consistent date shiftKeeps ages and intervals plausible
Free-text notesRedact or replaceFree text can contain anything

The data masking tutorial and pseudonymization vs anonymization go deeper. If realistic structure matters more than production correlations, consider synthetic data instead.

Try it

Draft a masking policy from column names only with the masking-policy starter among the free tools, then start free and apply it in the synthetic sandbox. The documentation describes policies in detail, and the integrations page lists the sources you can mask from today.