Use case
Healthcare software is judged on its exceptions. Paying an ordinary claim is the easy part; the defects that reach production are in denials and resubmissions, members with two coverages, labs flagged far outside their range, inpatient stays that span a month boundary, and refills that arrive a day early. A test dataset that lacks those cases lets the code that handles them go untested, and using real patient records to find them is exactly what a test environment should avoid.
The scenario
A QA lead for a payer platform wants to know, before writing a single test, which business situations a dataset actually contains. Counting is cheap and catches the common failure: a "realistic" dataset with no denied claims, or with lab values that never leave the normal range. The healthcare pack names its scenarios (routine care, claim denials, coordination of benefits, abnormal labs, pharmacy refills and more) and its synthetic bundle is generated from those models with a fixed seed.
The script below inventories the small public sample. Run it against a larger generated set or an export of a DataNivra-certified dataset and the same checks tell you whether the scenarios you need are present in useful numbers.
Run it
Python 3.10 or later, standard library only.
import csv
import io
import os
import urllib.request
from collections import Counter
BASE = os.environ.get("DATANIVRA_DOWNLOADS", "https://www.datanivra.com/downloads")
BUNDLE = f"{BASE}/packs/healthcare/2.1.0/synthetic"
def rows(name):
with urllib.request.urlopen(f"{BUNDLE}/{name}") as response:
return list(csv.DictReader(io.TextIOWrapper(response, encoding="utf-8")))
claims = rows("claims.claims.csv")
labs = rows("clinical.lab_results.csv")
encounters = rows("clinical.encounters.csv")
fills = rows("pharmacy.prescriptions.csv")
status = Counter(c["claim_status"] for c in claims)
print("claim status:", ", ".join(f"{k}={v}" for k, v in sorted(status.items())))
for c in claims:
if c["claim_status"] == "DENIED":
print(f"denied claim {c['claim_id']}: reason {c['denial_reason_code']}, paid {c['total_paid']}")
flags = Counter(lab["abnormal_flag"] for lab in labs)
print("lab flags:", ", ".join(f"{k}={v}" for k, v in sorted(flags.items())))
ranged = [lab for lab in labs if lab["result_value"] and lab["reference_low"] and lab["reference_high"]]
outside = [
lab
for lab in ranged
if not float(lab["reference_low"]) <= float(lab["result_value"]) <= float(lab["reference_high"])
]
print(f"lab values outside their reference range: {len(outside)} of {len(ranged)} with a range")
kinds = Counter(e["encounter_type"] for e in encounters)
print("encounter types:", ", ".join(f"{k}={v}" for k, v in sorted(kinds.items())))
refills = sorted({int(f["refill_number"]) for f in fills})
print(f"refill numbers present: {refills}")
print(
"synthetic rows only:",
all(r["_dn_provenance"] == "SYNTHETIC" for r in claims + labs + encounters + fills),
)Expected output
claim status: DENIED=1, PAID=9, PENDED=1
denied claim CLM00000006: reason ZR02, paid 0.00
lab flags: H=3, L=2, N=9
lab values outside their reference range: 5 of 13 with a range
encounter types: EMERGENCY=4, INPATIENT=1, OFFICE=3, TELEHEALTH=3
refill numbers present: [0, 1]
synthetic rows only: TrueThe sample holds one denied claim with a made-up reason code (ZR02; the pack's diagnosis and lab codes use the synthetic ZD… and ZL… prefixes in the same way) and a paid amount of zero, one pended claim, and five lab results flagged high or low. One lab result has no value at all — a missing result is a case your interface must handle too. Refills only reach number 1 at this scale; the pharmacy-refill scenario generates numbers up to 5 when you request it.
Schema
| Scenario (pack code) | Where to look | What to count |
|---|---|---|
HC_CLAIM_DENIALS | claims.claims | claim_status = DENIED, total_paid = 0, a denial reason |
HC_ABNORMAL_LABS | clinical.lab_results | abnormal_flag H/L/HH/LL, values outside the reference range |
HC_COORDINATION_OF_BENEFITS | enrollment.coverages | members with a primary and a secondary coverage |
HC_PHARMACY_REFILLS | pharmacy.prescriptions | refill_number 1–5, 30/90-day supplies |
HC_ROUTINE_CARE | all | office visits, paid claims, normal labs |
The bundle's QUICKSTART.md lists every scenario the pack defines, including negative-test scenarios with deliberately broken references.
What DataNivra does with your own data
With the pack activated, a dataset request can name scenarios from the pack: the agent subsets and masks your source inside your network and generates the requested scenario rows synthetically, tagging every row's provenance (SYNTHETIC, MASKED_PRODUCTION_LIKE, NEGATIVE_TEST). The dataset then passes the same certification gates as any other — masking completion, referential integrity, orphan detection, provenance — before it can be provisioned, and its manifest records provenance counts. The control plane only sees scenario codes, counts and gate outcomes. Certification here means DataNivra-certified against configured policy gates; it is not a regulatory or third-party certification and does not make a system compliant with any regulation.
Limits to plan around
- The public sample is generated at a small scale (a few members), so several scenarios appear once or not at all; generate at a larger scale for real coverage.
- Codes in the synthetic data (
ZD…diagnoses,ZL…lab tests,ZR02-style denial reasons) are deliberately not real code-set values. Tests that validate against a licensed code set need a mapping layer or masked real codes. - Scenario coverage shows that cases exist, not that your application handles them correctly.
Next steps
Read the healthcare test data tutorial, browse the Healthcare and Life Sciences pack, and plan the scenario mix with the synthetic scenario planner.
Synthetic downloads
Files from the Healthcare & Life Sciences pack 2.1.0 bundle. Everything in it is synthetic, generated from a fixed seed, and listed with its SHA-256 digest in the bundle's MANIFEST.json.
synthetic/claims.claims.csv synthetic/clinical.lab_results.csv synthetic/clinical.encounters.csv synthetic/pharmacy.prescriptions.csv schema/clinical.sql QUICKSTART.md
Industry pack: Healthcare & Life Sciences — its entities, scenarios, policy templates and the complete asset bundle.
Try it with DataNivra
The synthetic playground walks through discovery, subsetting, masking, validation and certification in your browser, with no account. Starting free gives you the hosted synthetic sandbox; your own sources need an agent in your network.