Tutorial
Test data ages. Schemas change, business rules change, reference data changes, and the dataset that was perfect in January starts producing confusing results by June. A Refresh rebuilds the dataset so it reflects the current world again.
Refresh cadence is the rhythm of those rebuilds: how often, triggered by what, and for which environments.
Text description
- Refresh (Your environment): Schedule or trigger → Rebuild in your environment → Re-certify → Swap into environment → Retire old version
Why it matters
Stale data creates false failures and false confidence. A test fails because a lookup table in QA lacks a new product code; another passes because the data never exercises a rule added last quarter. Schema drift is the sharpest case: a new column in production has no counterpart in test, so code touching it cannot be tested at all.
Refreshing too often has costs too. Every rebuild consumes compute and can disrupt teams in the middle of a test cycle, especially in UAT where users prepare scenarios by hand. The right cadence balances freshness against stability, and it differs across Environment tiers.
Example
A synthetic insurer (invented for this tutorial) sets these cadences:
| Environment | Cadence | Trigger |
|---|---|---|
| DEV | nightly | schedule |
| QA | weekly, Sunday night | schedule, plus on schema change |
| SIT | at the start of each integration cycle | release calendar |
| UAT | on request, frozen during sign-off | business owner |
| Performance | monthly | schedule |
When a new column appears in the claims system mid-week, the QA refresh is triggered early. The new column is classified before use, the dataset is rebuilt and re-certified, and it replaces the old version only after certification passes. UAT is left untouched until sign-off ends.
How DataNivra approaches it
Refresh policies are versioned and attached to dataset requests: a schedule, triggers and a retention rule for old versions. When a refresh is due, the control plane queues a declarative command; the agent leases it, rebuilds the dataset inside your environment, and runs the full Certification gates again. Only a certified version is swapped into the target environment; if certification fails, the previous certified version stays in place and the failure is reported.
Old versions are retired according to retention policy so they do not accumulate. Throughout, the control plane sees schedules, job states, counts and evidence references, never rows. Learn how drift is handled in schema drift and failure recovery.
Choosing a cadence: questions to ask
- How often does the schema change? If releases add columns every sprint, refresh at least once per release so new columns are classified and masked before code touches them.
- How long do test cycles run? Never replace a dataset in the middle of a UAT cycle; schedule the swap between cycles and keep the previous certified version available for reruns.
- What does a rebuild cost? Large estates may justify incremental refresh or smaller subsets rather than frequent full rebuilds. The storage and cost calculator among the free tools compares scenarios.
- Who is affected by a swap? Publish new versions alongside old ones and let teams move when ready, rather than overwriting in place.
- How many versions do you keep? Keeping every version forever multiplies storage and exposure; keeping only the latest makes old defects impossible to reproduce. A small retention window of certified versions, with older ones deprovisioned and removed, is usually the right balance.
- Does production volume change seasonally? Year-end, open enrolment or holiday peaks change what realistic data looks like; refresh before those periods so tests reflect them.
- What triggers an off-cycle refresh? A policy change, a new sensitive column found by discovery, or a failed certification should all trigger a rebuild.
Database-specific considerations are in the guides, and the integrations page lists supported sources. Scheduling and triggers are covered in the documentation. You can try the workflow on synthetic data when you start free.
Key takeaways
- Stale data produces both false failures and false passes.
- Set cadence per environment, not one rule for all.
- Re-certify on every refresh and swap in only certified versions.