Use case
A member identifier rarely lives in one table. In a payer system it is the primary key of the member table, a foreign key on every claim, a reference on lab results and prescriptions, and often a copy in a data warehouse and a CRM. If masking replaces it with a random value per table, every one of those joins breaks and the test data stops behaving like the system it came from. Deterministic masking solves this: the same input always produces the same replacement, so relationships survive while the original value does not.
The scenario
Your QA team needs claims test data in which a claim still belongs to the right member, but the member identifier in the test environment must not be the real one. The requirement has three parts: consistency (the same id masks to the same token in every table and every refresh), irreversibility without the key (nobody holding only the test data can compute the original), and separation (the token for a member id must not equal the token for, say, a subscriber id that happens to have the same value).
A keyed hash — HMAC-SHA256 — meets all three when the key is secret and the input is prefixed with a domain label. The example below shows the mechanics on synthetic data you can download, so you can see the joins hold before you trust the idea with anything else.
Run it
Python 3.10 or later, standard library only. The script downloads two synthetic CSV files from the healthcare pack bundle and masks the member id on both sides of the join.
import csv
import hashlib
import hmac
import io
import os
import urllib.request
BASE = os.environ.get("DATANIVRA_DOWNLOADS", "https://www.datanivra.com/downloads")
BUNDLE = f"{BASE}/packs/healthcare/2.1.0/synthetic"
# A demonstration key. In DataNivra the key lives in your own secret store and only its
# reference (vault://..., azure-kv://..., env://...) is ever configured.
KEY = os.environ.get("MASKING_KEY", "demo-key-public-not-secret").encode()
def rows(name):
with urllib.request.urlopen(f"{BUNDLE}/{name}") as response:
return list(csv.DictReader(io.TextIOWrapper(response, encoding="utf-8")))
def pseudonym(value, domain="member"):
digest = hmac.new(KEY, f"{domain}:{value}".encode(), hashlib.sha256).hexdigest()
return f"M-{digest[:12]}"
members = rows("enrollment.members.csv")
claims = rows("claims.claims.csv")
before = sum(1 for c in claims if c["member_ref"] in {m["member_id"] for m in members})
masked_ids = {pseudonym(m["member_id"]) for m in members}
masked_refs = [pseudonym(c["member_ref"]) for c in claims]
after = sum(1 for ref in masked_refs if ref in masked_ids)
print(f"members: {len(members)}, claims: {len(claims)}")
print(f"claims joined to a member before masking: {before}")
print(f"claims joined to a member after masking: {after}")
print(f"same input, same token: {pseudonym('MBR00000001') == pseudonym('MBR00000001')}")
print(f"MBR00000001 -> {pseudonym('MBR00000001')}")
print(
f"different domain, different token: {pseudonym('MBR00000001', 'subscriber') != pseudonym('MBR00000001')}"
)Expected output
members: 5, claims: 11
claims joined to a member before masking: 11
claims joined to a member after masking: 11
same input, same token: True
MBR00000001 -> M-c63a63983f79
different domain, different token: TrueRun it twice and the token for MBR00000001 is identical; set MASKING_KEY to any other value and every token changes while the join count stays at 11. That is the whole property in two lines: consistency comes from the key, secrecy comes from keeping the key out of the test environment.
Schema
| Table | Column | Role in the join | Expected classification |
|---|---|---|---|
enrollment.members | member_id | primary key | direct identifier |
claims.claims | member_ref | foreign key to members.member_id | direct identifier, health data |
claims.claim_lines | member_ref | denormalised copy of the same id | direct identifier, health data |
The portable DDL in the bundle (schema/enrollment.sql, schema/claims.sql) marks each sensitive column with the classification the pack expects; on your own estate the agent classifies the real schema itself.
What DataNivra does with your own data
The agent applies the same principle inside your network with the production key held in your secret store: keyed pseudonymization for identifiers, format-preserving strategies where an application validates formats, and shared masking domains so a member id masked in the claims database equals the one masked in the warehouse. The engine's tests check determinism, key separation and cross-system identity consistency; certification refuses a dataset in which a sensitive column was left unmasked. The control plane receives policy references and aggregate counts, never the key and never a value. Certification here means DataNivra-certified against configured policy gates; it is not a regulatory or third-party certification and does not make a system compliant with any regulation.
Limits to plan around
- Deterministic tokens are pseudonyms, not anonymisation: anyone with the key and a candidate value can recompute the token. Protect the key like the data it protects.
- Low-cardinality values (a gender code, a state) stay guessable even when hashed; use generalisation or synthetic replacement for those.
- The script truncates the digest to 12 hex characters for readability. The product sizes tokens to the column and checks for collisions.
- This page demonstrates the technique on synthetic files. It is not the production engine and it does not certify anything.
Next steps
Read the tutorial on cross-system deterministic masking, compare pseudonymization and anonymization, then try the full masking flow in the synthetic playground.
Synthetic downloads
Files from the Healthcare & Life Sciences pack 2.1.0 bundle. Everything in it is synthetic, generated from a fixed seed, and listed with its SHA-256 digest in the bundle's MANIFEST.json.
Industry pack: Healthcare & Life Sciences — its entities, scenarios, policy templates and the complete asset bundle.
Try it with DataNivra
The synthetic playground walks through discovery, subsetting, masking, validation and certification in your browser, with no account. Starting free gives you the hosted synthetic sandbox; your own sources need an agent in your network.