Guides

Test data management guides

Practical, original guides for engineers and data owners. Connector guides state plainly what works today and what is still coming next or on the roadmap.

Looking for the basics first? Start with the learning center. See every connector and its status on the integrations page.

Databases, warehouses and storage

How to build test data from each source, with its current connector status read from the product registry.

  • Test data management for PostgreSQL

    Available now

    How to build masked, referentially intact PostgreSQL test datasets inside your own network, with a read-only proof on the source and certified loads into QA.

  • Test data management for SQL Server

    Available now

    What test data management for Microsoft SQL Server involves, what DataNivra's SQL Server connector does today, its tested limits, and the export path for everything outside them.

  • Test data management for Oracle Database

    Available now

    Building safe Oracle test data - read-only extraction, entity subsets, consistent masking - and what DataNivra's Oracle connector covers today and what it does not.

  • Test data management for MySQL

    Available now

    How to get realistic, masked MySQL and MariaDB test data without dumping production into QA, and what the MySQL connector does and does not do yet.

  • Test data management for Snowflake

    Available now

    Safe test data from Snowflake - clones versus masked subsets, read-only roles, the status of DataNivra's Snowflake connector and an unload path for teams that prefer files.

  • Test data management for Databricks

    Available now

    Test data for Databricks lakehouses - Delta tables, Unity Catalog grants, masked subsets - and the current status of DataNivra's Databricks connector.

  • Test data management for Amazon S3

    Available now

    Build masked, certified test datasets from Parquet and CSV files in Amazon S3 or S3-compatible storage, with the agent running in your own AWS VPC.

  • Test data management for Azure Data Lake Storage

    Available now

    Build masked, certified test datasets from Parquet and CSV files in Azure Blob Storage or ADLS Gen2, with the DataNivra agent running inside your own Azure VNet.

  • Test data management for mainframes

    Preview

    Mainframe test data - EBCDIC records, COBOL copybooks, packed decimal and identifiers shared with modern systems - and what DataNivra's preview mainframe file connector does and does not do yet.

Industries

What makes test data hard in regulated industries, and how to handle it.

  • Insurance test data

    What makes insurance test data hard - policies, claims, premiums and long histories - and how to build it safely with the DataNivra Insurance pack.

  • Test data management for banking

    Realistic, safe test data for core banking, payments, cards and KYC/AML screening - valid formats, consistent identities across systems, and edge cases a random sample never contains.

  • Healthcare test data

    Why health data is the hardest test data to get right, and how to keep member, claim and encounter links intact without moving PHI.

  • Financial-services test data

    How to build realistic banking and card test data that keeps balances reconciling and card numbers valid, without exposing cardholder data.

Practices and concepts

The techniques behind safe, useful test data.

  • Masking vs tokenization for test data

    The practical difference between data masking and tokenization, when each fits test data, and how DataNivra implements both without sending values out of your network.

  • TDM in CI/CD

    How pipelines can request certified test data on demand, wait for it safely, and tear it down again — without ever handling raw rows.

  • Synthetic vs masked production-like data

    When to mask a governed subset of production, when to generate synthetic records, and why many teams combine both.

  • Refresh cadence

    How often test data should be rebuilt, what triggers a refresh, and how to swap in new versions without disrupting teams.

  • Referential integrity

    Why every foreign key in a test dataset must resolve, and how masking and subsetting can quietly break joins across tables and systems.

  • Customer-resident data processing

    What it means for every row-level operation to run inside your own network, and how a hosted service can still coordinate the work.

  • Dataset certification

    What it means to certify a test dataset, which gates it must pass, and why a failed gate must block provisioning rather than warn.