Learning center
Test data management glossary
Short, plain definitions of the terms used across the learning center.
- Agent
- The containerised DataNivra runtime installed in your network. It connects outbound only, leases work from the control plane and executes it locally.
- Anonymization
- Irreversibly removing the ability to identify individuals, directly or by combining attributes. Hard to achieve and hard to prove for rich datasets.
- Audit trail
- An append-only record of who did what and when. In DataNivra each event carries the hash of the previous one, so tampering is detectable.
- Business entity
- The real-world thing a test is about — a member, a customer, an order — together with all records that describe it across tables and systems.
- Cardholder data
- Payment card numbers and associated data whose handling is governed by payment-industry security standards.
- Certification
- Automated gates that prove a dataset meets policy — coverage, masking, referential integrity, schema, quality and provenance — before it may be provisioned. Masked is not certified.
- CI/CD
- Continuous integration and delivery: automated build, test and deployment pipelines that increasingly need on-demand test data.
- Classification
- Labelling each field with a sensitivity class (for example direct identifier, quasi-identifier, financial, health) so policy can decide how to treat it.
- Control plane
- The DataNivra-hosted service for people and policy: console, REST API, dataset requests, job state, approvals, schedules and audit. It stores metadata, never raw rows.
- Customer-resident processing
- Running all row-level work (profiling, subsetting, masking, synthesis, validation, provisioning) inside the customer’s own security boundary.
- Data plane
- The customer-resident part of DataNivra — agent, engine and industry packs — that performs every operation touching rows.
- Data product
- A dataset managed like a product: named owner, versioning, documentation, quality guarantees, consumers and a retirement plan.
- Deterministic masking
- Masking where the same input and key always produce the same output, so joins and cross-system links keep working.
- Egress guard
- The agent component that inspects every outbound message against allow-listed schemas and content detectors and blocks anything that looks like raw data.
- Environment tiers
- The chain of non-production environments — DEV, QA, SIT, UAT and performance — each with different data needs.
- Evidence reference
- A pointer (URI plus checksum) to evidence such as a certification report that remains stored in the customer environment.
- Fail closed
- When the privacy or integrity state is uncertain, stop and publish nothing rather than proceed and hope.
- Format-preserving substitution
- Replacing a value with one of the same format — length, character classes, check digits — so validation logic in the application still accepts it.
- Industry pack
- A versioned plugin with domain entities, relationships, detection and masking presets, synthetic scenarios and tutorials for one industry.
- Lease
- A time-bound grant that lets one agent execute one command. Expired leases return the work to the queue; commands execute at most once.
- Masking
- Transforming sensitive values into realistic substitutes so data stays useful for testing without exposing the original values.
- PHI
- Protected health information: health data linked to an identifiable individual, subject to strict handling rules in many jurisdictions.
- Platform integrity
- Assurance that the platform itself behaves as specified: correct versions, verified policies, at-most-once execution and tamper-evident records.
- Provisioning
- Delivering a certified dataset into a target environment such as DEV, QA, SIT, UAT or performance.
- Pseudonymization
- Replacing identifiers with consistent substitutes. Linkage is preserved, and re-identification is possible only with the key or lookup, which must be protected.
- Quasi-identifier
- An attribute that is not identifying on its own (postcode, birth date, gender) but can identify someone in combination with others.
- Referential closure
- The complete set of related records needed so that no reference in a subset dangles.
- Referential integrity
- The property that every reference (foreign key) points to a record that exists, within a database and across related systems.
- Refresh
- Rebuilding and re-certifying a dataset so an environment’s data stays current with schema and business changes.
- Schema drift
- A change in source structure — new, renamed, retyped or dropped columns or tables — relative to what policies were approved for.
- Secret reference
- The name or URI of a secret (for example a vault path), used in configuration instead of the secret value itself.
- Sensitive-data discovery
- Finding which tables and columns hold personal, health, financial or confidential information, using metadata, local profiling and detection rules.
- Subsetting
- Selecting a smaller, representative slice of data driven by business entities, while bringing along every related record those entities need.
- Synthetic data
- Data generated from rules, templates or models rather than copied from production. It contains no real records, though its realism depends on the generator.
- Test data management (TDM)
- The discipline of supplying the right data, in the right shape and privacy state, to every non-production environment when teams need it — and retiring it when they do not.
- Zero raw-production-data egress
- The guarantee that raw production values never leave the customer boundary for the DataNivra control plane. Copies may exist inside the boundary; none leave it.