Tutorial
These two words are often used as if they were interchangeable. They are not. Pseudonymization replaces identifiers with consistent substitutes, so records remain linkable and could be re-identified by someone holding the key or lookup. Anonymization aims to remove the possibility of identifying anyone at all, even by combining attributes.
Most useful test data is pseudonymized, and it is important to call it that.
Text description
- Pseudonymization: Original identifier → Keyed substitute → Consistent across systems → Reversible only with key
- Anonymization: Original identifier → Generalize / suppress / synthesize → No consistent link → Goal: not re-identifiable
Why it matters
Privacy regulations and internal policies generally treat pseudonymized data as still personal data, with obligations that continue to apply, while truly anonymized data may fall outside some of those obligations. Labelling pseudonymized data as "anonymous" therefore leads to it being handled more loosely than it should be.
The distinction also drives design. Pseudonymization preserves linkage — exactly what integration and regression testing need. Anonymization deliberately breaks or blurs linkage, which suits analytics sharing or demos but can make test data less realistic. Rich records with many quasi-identifiers are hard to anonymize convincingly without losing the detail tests depend on.
Example
The same synthetic customer treated two ways:
| Field | Original (synthetic) | Pseudonymized | Anonymized |
|---|---|---|---|
customer_id | C-1001 | C-7730 (same in every system) | removed |
birth_date | 1979-11-02 | 1979-12-14 (shifted) | 1975–1979 band |
postcode | AB1 2CD | AB9 4XY (substitute) | AB region only |
balance | 2,410.55 | 2,410.55 | 2,000–2,500 band |
The pseudonymized version still supports a test that follows C-7730 from CRM to billing. The anonymized version cannot, but it could be shared far more widely. All values are invented.
When neither is enough, fully synthetic records avoid the question by containing no real individuals, at the cost of depending on the generator for realism.
How DataNivra approaches it
DataNivra is explicit about what it produces. Deterministic keyed masking is pseudonymization: keys are held in your secret store, referenced by name, and never sent to the control plane. Generalisation and suppression rules are available where you need to reduce identifiability further, and synthetic generation is available where no source access is appropriate.
We do not describe masked data as anonymous, and we do not claim masking alone makes data non-personal. Certification records which techniques were applied to which classified columns, so reviewers can judge the result against their own policy. See Data masking and Synthetic vs masked data.
Key takeaways
- Pseudonymized data is still personal data; treat it that way.
- Linkage is a feature for testing and a risk for re-identification.
- Choose deliberately and record the choice as evidence.