Documentation

Capacity planner

How the capacity and cost estimate works, which assumptions it uses, and why its numbers are estimates rather than measurements.

Purpose

The capacity planner compares storage, processing and cost for a test-data estate under a ladder of scenarios: full copies per environment, a right-sized subset, compressed snapshots, a reusable certified dataset shared by all environments, a retention policy and incremental refresh. Each scenario keeps the previous one's choices and adds one lever.

It is available in the console (Capacity & Cost) and, as an educational calculator, in the public learning center. Every result is an estimate from the assumptions shown, not a measurement, benchmark or price quote. Prices are illustrative placeholders to replace with your own rates.

Where the numbers come from

The authoritative arithmetic is the control plane's POST /v1/capacity/estimate. The planner calls it once per scenario and converts bytes to cost with your prices. The console and learning center use an exact TypeScript port so the comparison can update as you type; shared golden vectors prove the port and the control plane return identical numbers. In the console, Estimate with the control plane re-checks your plan against the server and shows whether the results match.

Only assumptions are sent to the control plane: sizes, ratios, counts and prices. No data is involved.

Model

  • Subset size = estate size × subset ratio; compressed size = subset × compression ratio.
  • Storage per copy = compressed size × (1 + (versions kept − 1) × fraction rewritten per refresh). The first version is stored in full; later versions store what changed.
  • Copies = 1 when a certified dataset is shared, otherwise one per environment.
  • Processing per month = refreshes per month × subset size × fraction rewritten × copies.
  • Cost = bytes ÷ 10^9 × price per GB, rounded half-to-even to cents.

The full-copy baseline uses a subset ratio and compression ratio of 1, no sharing, the baseline number of versions and full rebuilds.

Assumptions

AssumptionDefaultNotes
Production estate size100 TBDecimal: 1 TB = 10^12 bytes (about 0.91 TiB). Hypothetical.
Rows in the estate400,000,000,000Only used for the subset row estimate.
Non-production environments6DEV, QA, SIT, UAT, PERF and similar.
Subset ratio0.05Share of the estate a referentially intact subset keeps.
Compression ratio0.3Depends heavily on your data; validate it.
Versions kept (full-copy baseline)3Copies kept today when cloning production.
Versions kept (retention policy)2Certified versions kept under a retention policy.
Refreshes per month4
Fraction rewritten per incremental refresh0.21 means a full rebuild.
Storage price0.023 per GB-monthIllustrative only.
Processing price0.01 per GB processedIllustrative only.

Limits

Byte and row counts are limited to 9 PB (9 × 10^15) so every implementation can represent them exactly. The model ignores network transfer, licensing, staff time and the compute used for masking or synthesis; add those separately if they matter to your decision.

← All documentation