Learning center

Healthcare test data

Why health data is the hardest test data to get right, and how to keep member, claim and encounter links intact without moving PHI.

Tutorial

Healthcare systems are built around a handful of entities — members or patients, their coverage, the encounters they have with providers, and the claims those encounters produce. Almost every one of those records is protected health information or can become identifying once combined with other fields.

Testing a claims engine, an eligibility service or a care-management application therefore needs data that behaves like real health data: plan changes mid-year, claims that are denied and resubmitted, members who see many providers. The challenge is getting that behaviour without exposing a single real patient.

A healthcare entity graph (illustrative)An illustrative healthcare model: a member or patient has coverage under a plan, encounters with providers, and claims that reference encounters, providers and coverage. Subsets and masking must keep these links intact.Healthcare & life sciencesMember / patientCoverage & planEncounterProviderClaim
A healthcare entity graph (illustrative). An illustrative healthcare model: a member or patient has coverage under a plan, encounters with providers, and claims that reference encounters, providers and coverage. Subsets and masking must keep these links intact.
Text description
  1. Healthcare & life sciences: Member / patient → Coverage & plan → Encounter → Provider → Claim

Why it matters

Health records are unusually rich. A date of birth, a postcode and a rare diagnosis code together can point to one person even after names are removed, so simple column-by-column scrubbing is rarely enough. At the same time, healthcare applications depend on precise timing: eligibility windows, service dates, filing deadlines. Masking that shifts one date but not the related ones breaks adjudication logic and produces false defects. And because claims reference encounters, providers and coverage, a subset that loses one of those links makes whole test scenarios impossible to run. Getting this wrong means either risky data in test environments or data that is too broken to test with.

Masking a synthetic recordA synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.Before (synthetic example)name: Jordan Blakemember_id: M-40021dob: 1984-03-17email:jordan.b@example.testAfter masking · Your environmentname: Riley Chenmember_id: M-91877dob: 1984-05-02email:user_7f3a@example.testApproved masking policy, keyed & deterministic
Masking a synthetic record. A synthetic example record with name Jordan Blake, member id M-40021, birth date 1984-03-17 and email jordan.b@example.test passes through a keyed, deterministic masking policy and becomes name Riley Chen, member id M-91877, birth date 1984-05-02 and email user_7f3a@example.test. All values are invented.
Text description
  1. Before (synthetic example): name: Jordan Blake → member_id: M-40021 → dob: 1984-03-17 → email: jordan.b@example.test

    Connection: Approved masking policy, keyed & deterministic

  2. After masking (Your environment): name: Riley Chen → member_id: M-91877 → dob: 1984-05-02 → email: user_7f3a@example.test

Example

The following records are entirely synthetic and invented for this tutorial.

EntityBefore (synthetic)After masking (synthetic)
MemberM-40021, Jordan Blake, born 1984-03-17M-91877, Riley Chen, born 1984-05-02
EncounterE-7730 on 2025-02-03 for M-40021E-7730 on 2025-04-19 for M-91877
ClaimC-5512 for E-7730, service 2025-02-03C-5512 for E-7730, service 2025-04-19

Notice three things. The member identifier is replaced consistently, so the encounter and claim still point to the same member. The dates are shifted by the same offset for this member, so the interval between birth and service is preserved. Free-text clinical notes, not shown, would be suppressed or replaced with synthetic text rather than masked word by word.

How DataNivra approaches it

The healthcare industry pack ships an entity model for members, coverage, encounters, providers and claims, detection rules for health-plan identifiers and clinical codes, and masking presets such as per-member date shifting and deterministic pseudonymous identifiers. You review those presets and approve them as your own versioned policy.

Everything runs through customer-resident processing: the agent discovers and profiles your health data locally, builds a referentially closed subset driven by members, applies the approved policy and runs certification gates before anything reaches a QA environment. Only column names, classes, counts and evidence references reach the control plane.

The pack supports organisations running HIPAA-related privacy programmes by keeping PHI inside their environment and producing certification evidence. It does not, by itself, make a system compliant. See the healthcare pack for details.

Common pitfalls

  • Masking names but keeping full dates. A date of birth plus a postal code and an admission date can single out a patient even when the name is gone. Shift dates per patient and generalise locations where tests allow it.
  • Forgetting free text. Clinical notes, referral letters and claim comments can contain names, diagnoses and addresses. Redact or replace them wholesale rather than hoping a pattern catches everything.
  • Breaking the link between clinical and billing systems. If the patient id is masked differently in the EHR extract and the claims extract, integration tests fail for reasons that have nothing to do with the code. Use one deterministic key for both.
  • Copying reference data you do not need. Provider directories and code tables are often safe, but staff and clinician tables usually are not.

Before connecting a clinical source, check its connector status on the integrations page, estimate the subset with the free tools, and read the guides for your database. You can rehearse the whole flow on synthetic data when you start free; the documentation covers agent installation.

Key takeaways

  • Treat combinations of fields, not just names, as identifying.
  • Shift related dates together so clinical and billing logic still works.
  • Keep members, encounters, providers and claims linked across every system.