Guides

Insurance test data

What makes insurance test data hard - policies, claims, premiums and long histories - and how to build it safely with the DataNivra Insurance pack.

Guide

Insurers test some of the most rule-heavy software in any industry: rating engines, underwriting workflows, claims adjudication, billing and reinsurance. Every one of them depends on long-lived, deeply connected records, and every one of them handles personal data, often including health information in life, disability and health lines. This guide explains what good insurance test data looks like and how DataNivra supports it today.

Status: Insurance pack installed

DataNivra's industry packs bundle domain rules, entity templates and synthetic scenarios. The installed packs are healthcare and life sciences, financial services, insurance and general enterprise, and every plan includes them. The Insurance pack adds entity templates for policyholders, policies, coverages, claims, claimants, claim and premium payments, agents and commissions; detection rules for policyholder and claimant identity, policy and claim numbers and health information in claims; linked masking templates that keep joins intact; and claims scenarios such as reopened claims, lapsed policies and catastrophe surges. It is also one of the synthetic estates of the hosted sandbox. Like every pack, it supports your own privacy programme; it does not by itself make any system compliant with insurance, health-information or privacy law. The Insurance pack page shows the details.

What makes insurance data different

  • Long histories. A life policy can be in force for decades. Endorsements, renewals, lapses and reinstatements create many versions of the same contract, and tests of renewal or reserving logic need those histories intact.
  • Many parties per record. A single policy links policyholders, insured persons, beneficiaries, payers, brokers and, for claims, claimants, witnesses and providers. Each is a person whose data needs protection, and they are connected in ways a naive subset breaks.
  • Claims carry the most sensitive content. Loss descriptions, adjuster notes and medical reports are free text that can contain anything.
  • Rating needs realistic distributions. Premium calculation depends on ages, locations, vehicle or property attributes and claim history. Random values produce premiums that no rating rule would ever produce, so tests pass without exercising anything.
  • Money must reconcile. Premiums billed, collected and refunded, and claim reserves and payments, must add up for finance tests to mean anything.

Building an entity-based subset

Pick the root entity your tests revolve around, usually the policyholder or the policy, then let the subset follow the relationships: policies to coverages and endorsements, policies to premium invoices and payments, policies to claims, claims to payments and reserves, and each record to its parties. Where the core system enforces relationships only in application code, declare them in policy so the subset stays referentially intact. Choose roots deliberately: a stratified selection across product lines and states, plus edge cases such as lapsed and reinstated policies, is far more useful than a random percentage. The subsetting tutorial explains the strategies.

Masking that keeps rating and claims logic working

  • Replace names, addresses, emails and phone numbers of every party with synthetic values.
  • Pseudonymise policy numbers, claim numbers and party ids deterministically under your key, so the same policy has the same pseudonym in the policy system, the billing system and the claims system.
  • Shift dates of birth and event dates by an entity-consistent offset, preserving ages at inception and the intervals between loss and report.
  • Redact or replace free-text claim notes; do not try to mask them word by word.
  • Keep amounts where finance tests need exact reconciliation, or perturb them consistently where amounts themselves are sensitive.

Health information in life and disability claims deserves the same treatment as clinical data; the healthcare test data tutorial covers it.

When synthetic data is the better answer

For new products, rating changes and high-volume performance tests, synthetic data is often better than masked data: you can generate portfolios that do not exist yet, boundary cases for every rating factor, and volumes beyond production. DataNivra's synthetic generator labels every row by scenario (normal, boundary, rare, high-volume and explicitly labelled negative-test data), so broken test data never masquerades as normal data. The synthetic-scenario planner among the free tools helps you plan the mix, and synthetic vs masked data explains the trade-off.

Sources

Insurance cores commonly run on databases whose DataNivra connectors are not shipped yet, such as Oracle, SQL Server or mainframe files. PostgreSQL, file sets, Amazon S3 and Azure Data Lake are available today; the others can be used through Parquet or CSV exports inside your network. Check the integrations page for current status.

Next steps

Start free with the synthetic sandbox to walk through discovery, subsetting, masking and certification, and read the documentation when you are ready to install an agent. The financial-services tutorial is a useful companion for premium billing and payments.