Data masking
How masking replaces sensitive values with realistic substitutes, why determinism matters, and how to keep masked data useful for testing.
Topics
Masking replaces sensitive values with realistic substitutes so test data keeps its shape and meaning without exposing people. Doing it well means finding every sensitive column first, choosing a technique that matches the risk, and keeping the same person recognisable across every system that holds them.
The pillar explains the technique; the cluster pages go deeper into classification, terminology, cross-system consistency and a starter policy you can adapt.
How masking replaces sensitive values with realistic substitutes, why determinism matters, and how to keep masked data useful for testing.
How to find sensitive fields across databases and files, classify them, and turn findings into approved policy without exposing values.
The practical difference between pseudonymized and anonymized data, why the distinction matters for test data, and how to choose.
The practical difference between data masking and tokenization, when each fits test data, and how DataNivra implements both without sending values out of your network.
How to mask the same customer consistently across databases, warehouses and files, so joins between systems still work in test without revealing the real identity.
Mask documents, chat and ticket exports inside your network and generate synthetic RAG evaluation sets - a Preview with rule-based detection and its limits stated.
A runnable example that pseudonymizes member ids in two synthetic healthcare tables with a keyed hash and proves every claim still joins to its member afterwards.