Tutorial
Full copies of production are expensive in ways that are easy to miss. Every environment holds its own copy, every copy needs storage, backups and licences, and every refresh moves the full volume again.
Optimising test data cost is mostly about not moving or keeping data nobody needs. The levers are simple; the discipline is applying them consistently.
Text description
- Levers: Subset, do not clone → Compressed formats → Reuse certified versions → Right-size compute → Expire unused datasets
Why it matters
Storage and compute for non-production environments can quietly rival production itself, especially when each team keeps its own clones. Large copies also slow everything down: refreshes take longer, so they happen less often, so data gets stale. And every extra copy is another place sensitive data could be exposed. Reducing volume therefore improves cost, speed and privacy at the same time. The goal is not the smallest possible dataset, but the smallest one that still exercises the behaviour each test needs.
Example
A synthetic logistics company (figures invented for illustration) reviews its QA estate:
| Before | After |
|---|---|
| A full clone per team, four teams | One shared certified subset per release, reused by all four |
| Monthly full refresh | Weekly refresh of the subset |
| Old copies kept "just in case" | Versions expire after two releases |
| Large fixed compute for every job | Job size chosen from the subset's volume |
The subset is driven by a few thousand orders and their related customers, shipments and invoices, plus synthetic edge cases for rare routes. Performance testing keeps a larger dataset, because volume is what that environment is for.
How DataNivra approaches it
The first lever is Subsetting: entity-driven subsets with referential closure mean you ship the slice you need instead of the whole estate. The second is reuse: certified versions are catalogued as a Data product, so several teams can consume one version rather than build their own. The third is right-sizing: jobs run on compute inside your environment, sized to the work, and the engine can use a distributed execution adapter for large volumes. The fourth is expiry: retention policies retire old versions automatically, and Refresh replaces rather than duplicates.
Where targets support it, compressed formats reduce footprint further. Usage is metered from aggregate metadata such as job durations and volumes, so you can see where the cost goes without any rows leaving your environment. See pricing for how metering works.
Key takeaways
- Subset instead of cloning, and reuse certified versions.
- Size compute to the job, and expire what nobody uses.
- Less data moved is cheaper, faster and safer.