Learning center

Learn test data management

Plain-language tutorials for engineers, testers, data owners and security reviewers. Each one explains why the topic matters, shows an example with synthetic data, and describes how DataNivra approaches it.

Prefer to listen? Try the narrated walkthroughs. Unsure about a term? See the glossary. Curious what test data costs to store and refresh? Try the capacity and cost planner.

  1. Tutorial 1

    What is Test Data Management?

    What test data management is, who needs it, and the lifecycle that turns "we need data" into a safe, certified dataset in a test environment.

  2. Tutorial 2

    Why production copies are risky

    Why cloning production into test environments creates privacy, security and operational risk, and what a governed alternative looks like.

  3. Tutorial 3

    DEV vs QA vs SIT vs UAT vs Performance

    How data needs differ across development, QA, system integration, user acceptance and performance environments, and how to serve each one.

  4. Tutorial 4

    Data discovery and classification

    How to find sensitive fields across databases and files, classify them, and turn findings into approved policy without exposing values.

  5. Tutorial 5

    Data masking

    How masking replaces sensitive values with realistic substitutes, why determinism matters, and how to keep masked data useful for testing.

  6. Tutorial 6

    Pseudonymization vs anonymization

    The practical difference between pseudonymized and anonymized data, why the distinction matters for test data, and how to choose.

  7. Tutorial 7

    Cross-system deterministic masking

    How to mask the same customer consistently across databases, warehouses and files, so joins between systems still work in test without revealing the real identity.

  8. Tutorial 8

    Data subsetting

    How to carve small, representative, referentially complete slices out of large databases, driven by the business entities your tests care about.

  9. Tutorial 9

    Referential integrity

    Why every foreign key in a test dataset must resolve, and how masking and subsetting can quietly break joins across tables and systems.

  10. Tutorial 10

    Cross-system test-data subsetting

    How to give QA one customer and every related policy, claim, account and transaction across several systems, without dangling references or unrelated records.

  11. Tutorial 11

    Synthetic vs masked production-like data

    When to mask a governed subset of production, when to generate synthetic records, and why many teams combine both.

  12. Tutorial 12

    Dataset certification

    What it means to certify a test dataset, which gates it must pass, and why a failed gate must block provisioning rather than warn.

  13. Tutorial 13

    Refresh cadence

    How often test data should be rebuilt, what triggers a refresh, and how to swap in new versions without disrupting teams.

  14. Tutorial 14

    Test data as a product

    Treating datasets like products with owners, versions, documentation and a retirement plan, instead of one-off copies nobody maintains.

  15. Tutorial 15

    Storage and compute optimization

    Practical levers for reducing the cost of test data — subsetting, reuse, right-sized jobs and expiry — without sacrificing realism.

  16. Tutorial 16

    TDM in CI/CD

    How pipelines can request certified test data on demand, wait for it safely, and tear it down again — without ever handling raw rows.

  17. Tutorial 17

    Healthcare test data

    Why health data is the hardest test data to get right, and how to keep member, claim and encounter links intact without moving PHI.

  18. Tutorial 18

    Financial-services test data

    How to build realistic banking and card test data that keeps balances reconciling and card numbers valid, without exposing cardholder data.

  19. Tutorial 19

    Customer-resident data processing

    What it means for every row-level operation to run inside your own network, and how a hosted service can still coordinate the work.

  20. Tutorial 20

    Zero raw-production-data egress

    The precise privacy promise behind DataNivra, how it is enforced in code on both sides of the boundary, and why it is not “zero copy”.

  21. Tutorial 21

    Platform integrity

    How a test data platform proves it behaves as configured — verified policies, at-most-once commands, revocable agents and fail-closed defaults.

  22. Tutorial 22

    Audit evidence

    What evidence a test data programme needs, why audit events should be tamper-evident, and why the evidence files should stay in your environment.

  23. Tutorial 23

    Schema drift and failure recovery

    How to handle source schema changes and failed jobs safely — detect drift, classify new fields first, block uncertified output and recover cleanly.