Learning center

Storage and compute optimization

Practical levers for reducing the cost of test data — subsetting, reuse, right-sized jobs and expiry — without sacrificing realism.

Tutorial

Full copies of production are expensive in ways that are easy to miss. Every environment holds its own copy, every copy needs storage, backups and licences, and every refresh moves the full volume again.

Optimising test data cost is mostly about not moving or keeping data nobody needs. The levers are simple; the discipline is applying them consistently.

Storage and compute leversFive levers reduce cost: subset instead of cloning, use compressed columnar formats where targets allow, reuse certified versions, right-size compute for each job, and expire datasets nobody uses.LeversSubset, do notcloneCompressedformatsReuse certifiedversionsRight-sizecomputeExpire unuseddatasets
Storage and compute levers. Five levers reduce cost: subset instead of cloning, use compressed columnar formats where targets allow, reuse certified versions, right-size compute for each job, and expire datasets nobody uses.
Text description
  1. Levers: Subset, do not clone → Compressed formats → Reuse certified versions → Right-size compute → Expire unused datasets

Why it matters

Storage and compute for non-production environments can quietly rival production itself, especially when each team keeps its own clones. Large copies also slow everything down: refreshes take longer, so they happen less often, so data gets stale. And every extra copy is another place sensitive data could be exposed. Reducing volume therefore improves cost, speed and privacy at the same time. The goal is not the smallest possible dataset, but the smallest one that still exercises the behaviour each test needs.

Example

A synthetic logistics company (figures invented for illustration) reviews its QA estate:

BeforeAfter
A full clone per team, four teamsOne shared certified subset per release, reused by all four
Monthly full refreshWeekly refresh of the subset
Old copies kept "just in case"Versions expire after two releases
Large fixed compute for every jobJob size chosen from the subset's volume

The subset is driven by a few thousand orders and their related customers, shipments and invoices, plus synthetic edge cases for rare routes. Performance testing keeps a larger dataset, because volume is what that environment is for.

How DataNivra approaches it

The first lever is Subsetting: entity-driven subsets with referential closure mean you ship the slice you need instead of the whole estate. The second is reuse: certified versions are catalogued as a Data product, so several teams can consume one version rather than build their own. The third is right-sizing: jobs run on compute inside your environment, sized to the work, and the engine can use a distributed execution adapter for large volumes. The fourth is expiry: retention policies retire old versions automatically, and Refresh replaces rather than duplicates.

Where targets support it, compressed formats reduce footprint further. Usage is metered from aggregate metadata such as job durations and volumes, so you can see where the cost goes without any rows leaving your environment. See pricing for how metering works.

Key takeaways

  • Subset instead of cloning, and reuse certified versions.
  • Size compute to the job, and expire what nobody uses.
  • Less data moved is cheaper, faster and safer.