Learning center

Cross-system deterministic masking

How to mask the same customer consistently across databases, warehouses and files, so joins between systems still work in test without revealing the real identity.

Tutorial

Most business processes cross system boundaries. A customer signs up in a CRM, is billed from an ERP, appears in a claims database and ends up in a warehouse table used for reporting. Each system has its own copy of the customer's identifiers. Testing the process end to end needs every one of those copies masked — and masked the same way, or the systems stop agreeing about who the customer is.

Deterministic masking is the technique that makes this possible: the same input, masked with the same key, always produces the same output, wherever the masking runs.

Masking a synthetic recordA synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.Before (synthetic example)name: Jordan Blakemember_id: M-40021dob: 1984-03-17email:jordan.b@example.testAfter masking · Your environmentname: Riley Chenmember_id: M-91877dob: 1984-05-02email:user_7f3a@example.testApproved masking policy, keyed & deterministic
Masking a synthetic record. A synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.
Text description
  1. Before (synthetic example): name: Jordan Blake → member_id: M-40021 → dob: 1984-03-17 → email: jordan.b@example.test

    Connection: Approved masking policy, keyed & deterministic

  2. After masking (Your environment): name: Riley Chen → member_id: M-91877 → dob: 1984-05-02 → email: user_7f3a@example.test

Why it matters

When each system is masked on its own, with random replacement values, every join between systems breaks. The CRM says customer C-1041 became C-8812; the billing system, masked separately, turned the same C-1041 into C-2270. An integration test that follows a customer from sign-up to invoice now finds nobody, and teams quietly fall back to unmasked copies "just for this test".

Cross-system consistency is also a Referential integrity problem that no single database can enforce. The foreign key between a CRM table and a billing table does not exist in either catalog; it exists only in the business process. So the masking has to preserve it by construction, not by a database constraint.

There is a privacy side too. Consistent pseudonyms are still pseudonyms, not anonymous data: anyone with the key could recompute them. That is why the key must stay with the data owner, be versioned, and be rotated deliberately.

Referential integrity across tables and systemsParent records such as customers and accounts must exist for every child record such as transactions, cards and statements. When a key is masked, the same masked value must be used everywhere it appears so that joins still work across tables and systems.Parentscustomer (key C-1001)account (key A-2001 →C-1001)Childrentransaction → A-2001card → A-2001statement → A-2001Every child key resolves to a parent; masked keys stay consistent
Referential integrity across tables and systems. Parent records such as customers and accounts must exist for every child record such as transactions, cards and statements. When a key is masked, the same masked value must be used everywhere it appears so that joins still work across tables and systems.
Text description
  1. Parents: customer (key C-1001) → account (key A-2001 → C-1001)

    Connection: Every child key resolves to a parent; masked keys stay consistent

  2. Children: transaction → A-2001 → card → A-2001 → statement → A-2001

Example

A bank tests a new loan-servicing flow that touches three systems:

SystemIdentifier columnExample value (synthetic)
Customer database (PostgreSQL)customer_id100427
Loan servicing (SQL Server)borrower_ref100427
Reporting warehousecust_key'100427' (stored as text)

With keyed, deterministic masking all three values become the same pseudonym, for example 738215, because the same key and the same masking rule are applied to the same canonical value. The text and integer forms are compared canonically, so the warehouse's string key still matches.

Now consider a harder case: the loan system stores the customer's e-mail address, and the customer database stores it too, in different letter case. Normalising the value before masking (trimming and lower-casing an e-mail address) is what makes Ana.Lee@example.test and ana.lee@example.test land on the same masked address. Normalisation rules belong to the masking policy, versioned with it, so a later run reproduces the same results.

The last case is identifiers that differ between systems: the CRM calls the customer customer_id = 100427 while a legacy system calls the same person party_key = P-00093. Deterministic masking alone cannot link these, because the inputs are different. They need an explicit identity mapping that an owner approves, and any matching that looks at values must run where the data lives.

How DataNivra approaches it

Masking runs in the agent inside your network. Pseudonyms are computed with keyed HMAC derivations and format-preserving permutations. Format-preserving strategies keep the shape of the original (length, character classes, check digits where a strategy supports them), and because they are keyed permutations, two different inputs cannot map to the same pseudonym. The key never leaves your environment: masking rules carry only a key reference and a key version, which the agent resolves from your secret store.

Because the derivation depends on the key, the key version and the masking strategy — not on the table or column name — the same identifier masked in two systems with the same rule produces the same pseudonym. The engine's tests mask a CRM and a billing system in separate runs and verify that the masked keys still join, and the healthcare pack verifies consistent member, provider and pharmacy identifiers across four systems.

Rotating a key is a deliberate event: new datasets are masked with the new version, and pseudonyms from different key versions are intentionally unrelated, so old and new datasets do not link. Versioned semantic masking domains — one rule per business identity such as a person, an account or a policy, applied across differently formatted columns — and owner-approved identity mappings between different identifiers are in active development. Any value-assisted matching will run customer-resident; raw values are never sent to DataNivra Cloud.

To see keyed masking on synthetic data, open the interactive demo. Read data masking for the strategies themselves, cross-system test-data subsetting for selecting related records across systems, and the integrations page for which sources can be read today.

Key takeaways

  • Consistent pseudonyms across systems need one key, one rule and one normalisation per identity.
  • Deterministic does not mean anonymous: keep the key with the data owner and version it.
  • Different identifiers for the same entity need an approved identity mapping, not guessing.