Learning center

Capacity and cost planner for test data

Explore how test-data design choices change storage, processing and cost for a hypothetical 100 TB estate. Every value is an adjustable assumption and every result is an estimate.

Estimate — not a measurement or quote. Estimate from the assumptions shown — not a measurement, benchmark or price quote. Prices are illustrative list-price assumptions; replace them with your own rates.

Your assumptions

Sizes are decimal (1 TB = 10^12 bytes, so 100 TB is about 90.9 TiB). The calculator runs in your browser; nothing you type is sent anywhere.

Total size of the source data a full copy would duplicate. Default: a hypothetical 100 TB.
Only used to estimate how many rows a subset keeps.
DEV, QA, SIT, UAT, PERF, ... that each need the data.
Share of the estate a right-sized, referentially intact subset keeps.
Size of a compressed snapshot relative to the uncompressed subset.
How many historical copies teams keep today when they clone production.
How many certified versions a retention policy keeps.
How often the test data is refreshed.
Share of the data an incremental refresh rewrites (1 = full rebuild).
Illustrative storage list price; replace with your own rate.
Illustrative compute cost per GB rewritten by refreshes; replace with your own rate.
A label only — no conversion is applied.

Scenario comparison

Estimate: with every lever applied, about USD 81.40 per month, compared with USD 65,400.00 for full copies (99.9% lower).

Estimated monthly storage, processing and cost per scenario. Each row adds one lever to the row above. Estimates only.
ScenarioStorage footprintProcessed / monthTotal cost / monthvs full copiesRelative cost
Full copy per environment1.8 PB2.4 PBUSD 65,400.00baseline100% of the highest
Right-sized subset90 TB120 TBUSD 3,270.0095.0% lower5% of the highest
Compressed snapshots27 TB120 TBUSD 1,821.0097.2% lower3% of the highest
Reusable certified dataset4.5 TB20 TBUSD 303.5099.5% lower<1% of the highest
Retention policy3 TB20 TBUSD 269.0099.6% lower<1% of the highest
Incremental refresh1.8 TB4 TBUSD 81.4099.9% lower<1% of the highest

What each lever does

Full copy per environment
Cloning production into every test environment is the simplest approach, and the most expensive: each environment stores the whole estate, often several times over, and every refresh rewrites all of it. Model: Baseline: every environment holds an uncompressed full copy, kept for the baseline number of versions and fully rebuilt on each refresh.
Right-sized subset
Most tests need realistic, connected records, not every record. A subset keeps a slice of root entities plus everything they reference, so relationships still hold while the volume shrinks. Model: Keep only a referentially intact subset of the estate.
Compressed snapshots
Test datasets are written once and read many times, which suits compressed, columnar snapshot formats. The compression ratio depends heavily on your data, so treat it as an assumption to validate. Model: Store the subset as a compressed snapshot.
Reusable certified dataset
When a dataset has passed its privacy and quality checks, it can be published as an immutable, certified version. Environments then reuse that one version instead of each keeping a private copy. Model: Certify one immutable snapshot and let every environment reuse it.
Retention policy
Old versions are rarely needed for long. A retention policy keeps a small number of certified versions for reproducibility and removes the rest. Model: Keep only the configured number of certified versions.
Incremental refresh
If only part of the data changes between refreshes, rewriting just that part reduces both processing and the extra storage each new version needs. Model: Refresh only the changed fraction instead of rebuilding.

How the estimate is calculated

  • Subset size = estate size × subset ratio; compressed size = subset × compression.
  • Storage per copy = compressed size × (1 + (versions kept − 1) × fraction rewritten per refresh): the first version is stored in full, later versions store what changed.
  • Copies = 1 when a certified dataset is shared, otherwise one per environment. Processing per month = refreshes × subset size × fraction rewritten × copies.
  • Cost = size in GB (10^9 bytes) × your price, rounded to cents. The same model powers the capacity estimate in the DataNivra console.

Want the background? Read the learning center tutorials or the capacity planner documentation.