Learning center
Capacity and cost planner for test data
Explore how test-data design choices change storage, processing and cost for a hypothetical 100 TB estate. Every value is an adjustable assumption and every result is an estimate.
Estimate — not a measurement or quote. Estimate from the assumptions shown — not a measurement, benchmark or price quote. Prices are illustrative list-price assumptions; replace them with your own rates.
Your assumptions
Sizes are decimal (1 TB = 10^12 bytes, so 100 TB is about 90.9 TiB). The calculator runs in your browser; nothing you type is sent anywhere.
Scenario comparison
Estimate: with every lever applied, about USD 81.40 per month, compared with USD 65,400.00 for full copies (99.9% lower).
| Scenario | Storage footprint | Processed / month | Total cost / month | vs full copies | Relative cost |
|---|---|---|---|---|---|
| Full copy per environment | 1.8 PB | 2.4 PB | USD 65,400.00 | baseline | 100% of the highest |
| Right-sized subset | 90 TB | 120 TB | USD 3,270.00 | 95.0% lower | 5% of the highest |
| Compressed snapshots | 27 TB | 120 TB | USD 1,821.00 | 97.2% lower | 3% of the highest |
| Reusable certified dataset | 4.5 TB | 20 TB | USD 303.50 | 99.5% lower | <1% of the highest |
| Retention policy | 3 TB | 20 TB | USD 269.00 | 99.6% lower | <1% of the highest |
| Incremental refresh | 1.8 TB | 4 TB | USD 81.40 | 99.9% lower | <1% of the highest |
What each lever does
- Full copy per environment
- Cloning production into every test environment is the simplest approach, and the most expensive: each environment stores the whole estate, often several times over, and every refresh rewrites all of it. Model: Baseline: every environment holds an uncompressed full copy, kept for the baseline number of versions and fully rebuilt on each refresh.
- Right-sized subset
- Most tests need realistic, connected records, not every record. A subset keeps a slice of root entities plus everything they reference, so relationships still hold while the volume shrinks. Model: Keep only a referentially intact subset of the estate.
- Compressed snapshots
- Test datasets are written once and read many times, which suits compressed, columnar snapshot formats. The compression ratio depends heavily on your data, so treat it as an assumption to validate. Model: Store the subset as a compressed snapshot.
- Reusable certified dataset
- When a dataset has passed its privacy and quality checks, it can be published as an immutable, certified version. Environments then reuse that one version instead of each keeping a private copy. Model: Certify one immutable snapshot and let every environment reuse it.
- Retention policy
- Old versions are rarely needed for long. A retention policy keeps a small number of certified versions for reproducibility and removes the rest. Model: Keep only the configured number of certified versions.
- Incremental refresh
- If only part of the data changes between refreshes, rewriting just that part reduces both processing and the extra storage each new version needs. Model: Refresh only the changed fraction instead of rebuilding.
How the estimate is calculated
- Subset size = estate size × subset ratio; compressed size = subset × compression.
- Storage per copy = compressed size × (1 + (versions kept − 1) × fraction rewritten per refresh): the first version is stored in full, later versions store what changed.
- Copies = 1 when a certified dataset is shared, otherwise one per environment. Processing per month = refreshes × subset size × fraction rewritten × copies.
- Cost = size in GB (10^9 bytes) × your price, rounded to cents. The same model powers the capacity estimate in the DataNivra console.
Want the background? Read the learning center tutorials or the capacity planner documentation.