Learning center

Test data as a product

Treating datasets like products with owners, versions, documentation and a retirement plan, instead of one-off copies nobody maintains.

Tutorial

In many organisations, test data is a favour. Someone with database access makes a copy, hands it over, and moves on. Months later nobody knows where it came from, what is in it or whether it is still safe to use.

Treating test data as a Data product changes that. A dataset gets a purpose, an owner, a version history, documentation and consumers, and it is retired deliberately when it is no longer needed.

Test data as a productA dataset product has a lifecycle: a request describing entity, scope and environment; policy approval; build and certification; a catalogued version; consumption by teams; and expiry or retirement.Dataset productRequestPolicyapprovalBuild &certifyCataloguedversionConsumeExpire /retire
Test data as a product. A dataset product has a lifecycle: a request describing entity, scope and environment; policy approval; build and certification; a catalogued version; consumption by teams; and expiry or retirement.
Text description
  1. Dataset product: Request → Policy approval → Build & certify → Catalogued version → Consume → Expire / retire

Why it matters

Products are maintained; favours are not. When a dataset has an owner, someone notices when it breaks. When it has versions, teams can pin a known-good state for a release and compare results across versions. When it has documentation, new testers can find the scenario they need rather than building another copy. And when it has a retirement plan, old copies stop piling up, which reduces both storage cost and privacy exposure. The product view also makes demand visible: which datasets are used, by whom, and which ones nobody has touched in months.

Example

A synthetic payments team (all names invented) publishes a dataset product:

  • Name: "Cards QA — disputed transactions"
  • Owner: the QA lead for the payments squad
  • Contents: 1,200 synthetic-and-masked customers with card disputes, chargebacks and reversals
  • Environments: QA and SIT
  • Versions: v14 (current, certified), v13 (retained for the ongoing release), older versions retired
  • Refresh: weekly
  • Expiry: retire when the dispute-handling project closes

A tester needing chargeback scenarios finds this product in the catalogue instead of asking for a fresh copy, and a release team pins v13 until its regression cycle ends.

How DataNivra approaches it

Every dataset in DataNivra begins as a request that names the Business entity, scope, masking policy, target environment, refresh and retention. Requests go through policy approval, then the agent builds and certifies the dataset inside your environment. Each certified build becomes a catalogued version with metadata, aggregate metrics and evidence references visible in the console.

Consumers see what exists, what it contains in terms of entities and scenarios, and which version is current. Refresh policies keep versions current and retention policies expire them. Because the catalogue holds metadata only, it can be shared broadly without widening access to the data itself. See the product overview and the interactive demo.

Key takeaways

  • Give datasets owners, versions and documentation.
  • Make reuse easier than making another copy.
  • Plan retirement from the start.