Guide
"Masking" and "tokenization" are often used as if they meant the same thing. They overlap, but they solve different problems, and choosing the wrong one for test data leads either to useless datasets or to datasets that are less protected than they look. This guide explains the difference in practical terms and shows how DataNivra implements each.
Definitions that actually help
Masking replaces a sensitive value with a different value that is fit for purpose. The replacement might be a realistic synthetic name, a shifted date, a perturbed amount, a redaction marker or a keyed pseudonym. The defining property is that the original is not recoverable from the masked dataset, and the replacement is chosen to keep the data useful: formats, lengths, distributions and joins.
Tokenization replaces a value with a token, an opaque surrogate such as tok_..., and typically keeps a mapping so that authorised systems can relate tokens back to the original. Tokenization was designed for production: a payment system stores tokens instead of card numbers, and only a hardened vault can map them back. The defining property is the vault, and the vault is the most sensitive system in the design.
What each means for test data
Test environments need data that behaves like production. That usually means:
- values that pass validation (an email that looks like an email, a postal code in the right format);
- identifiers that are unique and join consistently across tables and systems;
- realistic spreads of dates and amounts, so business rules fire.
Masking is built for this. Tokens, by contrast, are opaque: a tokenized email fails validation, a tokenized date is not a date, and tests that parse or display these fields break. Tokenization is useful in test data for fields whose only job is to be a stable, unique reference, such as an external account reference that tests only compare for equality.
The bigger concern is reversibility. A reversible token vault reachable from test environments turns every test dataset back into production data for anyone who can call the vault. For test data, prefer irreversible transformations, and if you must keep a mapping, keep it inside the production boundary and never in lower environments.
How DataNivra implements masking
Masking policies assign a strategy to each classified column. Strategies include synthetic replacements for names, emails, phone numbers and addresses; date shifting that uses the same offset for the same entity so intervals stay realistic; numeric perturbation; redaction and nulling; hashing; keyed HMAC pseudonymisation; and a format-preserving permutation for bounded integer domains that is collision-free, so a masked primary key remains unique. Deterministic strategies use a key held in your secret store, which means the same customer id becomes the same pseudonym in every table, every system and every refresh, and referential integrity survives masking.
All masking runs in the agent inside your network. The key is referenced by name, never sent to DataNivra Cloud. Certification then checks that every column the policy requires to be masked actually was, and blocks the dataset otherwise.
How DataNivra implements tokenization
The TOKENIZE strategy uses a pluggable token vault. Two ship with the engine:
- a stateless HMAC vault that produces
tok_plus a keyed digest: deterministic, consistent across jobs that share the key, irreversible, with nothing stored; - a customer-local SQLite vault that issues random tokens and persists the mapping so tokens stay stable across runs. The mapping is keyed by a keyed fingerprint of the value, not the value itself, so the vault file holds no plaintext source values.
Customers can register an adapter to an existing enterprise tokenization service through the plugin registry. Any such vault stays inside your boundary.
Choosing between them
| Field | Good choice | Why |
|---|---|---|
| Customer name, email, phone | Synthetic replacement | Tests need valid, realistic formats |
| Customer or account id used in joins | Deterministic pseudonym or format-preserving permutation | Unique, consistent across tables |
| External reference compared only for equality | Token | Opaque is fine; stable across runs |
| Date of birth, admission date | Entity-consistent date shift | Keeps ages and intervals plausible |
| Free-text notes | Redact or replace | Free text can contain anything |
The data masking tutorial and pseudonymization vs anonymization go deeper. If realistic structure matters more than production correlations, consider synthetic data instead.
Try it
Draft a masking policy from column names only with the masking-policy starter among the free tools, then start free and apply it in the synthetic sandbox. The documentation describes policies in detail, and the integrations page lists the sources you can mask from today.