Use case
Test data goes stale in two directions. The schema moves on — a new column, a renamed table — and tests start failing for reasons unrelated to the change being tested. Or the data itself ages: dates drift into the past, statuses that were "open" in the snapshot are long closed in the application's logic, and the edge cases a release needs are not there. A refresh cadence is the agreement about how often a test environment gets a new certified version, and what is allowed to trigger one.
The scenario
A claims team runs integration tests every night and a full regression every week. Rebuilding test data nightly would waste agent compute and produce versions nobody tested against; never rebuilding means the weekly regression keeps running against a dataset from last quarter. The team wants a weekly scheduled refresh in a quiet hour, the option to request one on demand when a schema change lands, a ceiling on how many builds can run in a day, and no rebuild at all when the source has not changed since the last certified version.
DataNivra's industry packs ship that agreement as a policy template. The weekly template from the healthcare pack is a small JSON document you can read, review and adapt before a policy is ever approved.
Run it
Python 3.10 or later, standard library only. The script downloads the published template and computes the next four scheduled windows after a fixed date, so the output is reproducible.
import json
import os
import urllib.request
from datetime import datetime, timedelta, timezone
BASE = os.environ.get("DATANIVRA_DOWNLOADS", "https://www.datanivra.com/downloads")
URL = f"{BASE}/packs/healthcare/2.1.0/policies/HC_REFRESH_WEEKLY.json"
with urllib.request.urlopen(URL) as response:
policy = json.load(response)
print("cadence:", policy["cadence"], policy["timezone"])
print("skip when the source is unchanged:", policy["skip_if_source_unchanged"])
print("at most", policy["max_refreshes_per_day"], "refreshes a day; on demand:", policy["on_demand_allowed"])
# The cadence is a five-field cron expression: minute hour day-of-month month day-of-week.
minute, hour, dom, month, dow = policy["cadence"].split()
assert (dom, month) == ("*", "*"), "this sketch handles weekly/daily cadences only"
weekdays = range(7) if dow == "*" else [int(d) % 7 for d in dow.split(",")]
def next_runs(after, count):
runs, day = [], after.replace(hour=0, minute=0, second=0, microsecond=0)
while len(runs) < count:
at = day.replace(hour=int(hour), minute=int(minute))
if (at.isoweekday() % 7) in weekdays and at > after:
runs.append(at)
day += timedelta(days=1)
return runs
start = datetime(2026, 10, 9, 12, 0, tzinfo=timezone.utc)
for run in next_runs(start, 4):
print("next refresh:", run.strftime("%a %Y-%m-%d %H:%M UTC"))Expected output
cadence: 0 3 * * 1 UTC
skip when the source is unchanged: True
at most 4 refreshes a day; on demand: True
next refresh: Mon 2026-10-12 03:00 UTC
next refresh: Mon 2026-10-19 03:00 UTC
next refresh: Mon 2026-10-26 03:00 UTC
next refresh: Mon 2026-11-02 03:00 UTCMonday 03:00 UTC is before most teams start work in Europe and after most finish in the Americas, which is why the template uses it; move it to match your regression schedule.
Schema
| Field | Meaning |
|---|---|
cadence | cron expression for scheduled refreshes |
timezone | the zone the cron expression is evaluated in |
skip_if_source_unchanged | no new version when nothing changed since the last certified one |
max_refreshes_per_day | ceiling on refresh jobs of a dataset in a rolling 24-hour window |
on_demand_allowed | whether a person or pipeline may request an extra refresh |
incremental | rebuild only what changed (off in this template) |
What DataNivra does with your own data
An approved refresh policy schedules dataset builds for the agent. Each build produces a new immutable dataset version that must pass validation and certification before it can be provisioned; the previous certified version stays available, so a team can pin the version its tests last passed against. Before a build the engine compares schema fingerprints and table contents with the previous version: unchanged sources are skipped (reason SOURCE_UNCHANGED), and schema drift always forces a full rebuild rather than an incremental merge, so the new version is validated and certified from scratch. A request over the daily ceiling is refused with REFRESH_RATE_LIMITED; an on-demand request when the policy forbids it is refused with ON_DEMAND_REFRESH_DISABLED. Certification here means DataNivra-certified against configured policy gates; it is not a regulatory or third-party certification and does not make a system compliant with any regulation.
Limits to plan around
- The script evaluates only the simple weekly or daily cron shapes the pack templates use; it is a planning aid, not the product's scheduler.
- A refresh that is skipped because the source is unchanged produces no new version; tests that need fresh relative dates should use date shifting in the masking policy rather than more frequent builds.
- The ceiling counts every refresh job of the dataset in the last 24 hours, whatever its outcome. A burst of on-demand requests can use it up before the scheduled run.
Next steps
Read the refresh cadence tutorial and schema drift and failure recovery, then size your environments with the capacity planner.
Synthetic downloads
Files from the Healthcare & Life Sciences pack 2.1.0 bundle. Everything in it is synthetic, generated from a fixed seed, and listed with its SHA-256 digest in the bundle's MANIFEST.json.
Industry pack: Healthcare & Life Sciences — its entities, scenarios, policy templates and the complete asset bundle.
Try it with DataNivra
The synthetic playground walks through discovery, subsetting, masking, validation and certification in your browser, with no account. Starting free gives you the hosted synthetic sandbox; your own sources need an agent in your network.