Use case
Fixtures are the part of a test suite nobody owns. They start as a few rows typed into a JSON file, get copied between repositories, and slowly diverge from the schema they were meant to represent. When a team finally replaces them with something realistic, the replacement is often an export from a shared database — and a pull request that adds a CSV of customer records to a public repository is one review away from becoming an incident. Synthetic fixtures fetched at build time avoid both problems, as long as the pipeline can prove it received exactly the files that were published.
The scenario
A team wants every pull request to run integration tests against a small, realistic, relational dataset: members, coverages, claims and claim lines that join correctly. Nothing sensitive may be committed, the job must work on GitHub-hosted runners without any account or secret, and a tampered or truncated download must stop the job instead of producing confusing test failures.
Every DataNivra pack bundle publishes a MANIFEST.json with the SHA-256 digest of each file. The job below downloads the fixture files, refuses any file whose digest differs, loads them into an in-memory SQLite database and checks the relationships the tests rely on.
Run it
Save the script as ci/fetch_fixtures.py. Python 3.10 or later, standard library only.
import csv
import hashlib
import io
import json
import os
import sqlite3
import urllib.request
BASE = os.environ.get("DATANIVRA_DOWNLOADS", "https://www.datanivra.com/downloads")
PACK = f"{BASE}/packs/healthcare/2.1.0"
TABLES = ["enrollment.members", "enrollment.coverages", "claims.claims", "claims.claim_lines"]
def fetch(path):
with urllib.request.urlopen(f"{PACK}/{path}") as response:
return response.read()
# 1. Integrity first: every file must match the SHA-256 published in the bundle manifest.
manifest = {f["path"]: f["sha256"] for f in json.loads(fetch("MANIFEST.json"))["files"]}
db = sqlite3.connect(":memory:")
for table in TABLES:
path = f"synthetic/{table}.csv"
data = fetch(path)
if hashlib.sha256(data).hexdigest() != manifest[path]:
raise SystemExit(f"checksum mismatch for {path}: refusing to use it")
rows = list(csv.reader(io.StringIO(data.decode("utf-8"))))
name = table.replace(".", "_")
db.execute(f"CREATE TABLE {name} ({', '.join(rows[0])})")
db.executemany(f"INSERT INTO {name} VALUES ({', '.join('?' * len(rows[0]))})", rows[1:])
print(f"verified and loaded {table}: {len(rows) - 1} rows")
# 2. The checks your tests rely on: relationships hold and every row is synthetic.
orphans = db.execute(
"SELECT COUNT(*) FROM claims_claims c "
"LEFT JOIN enrollment_members m ON m.member_id = c.member_ref WHERE m.member_id IS NULL"
).fetchone()[0]
lines = db.execute(
"SELECT COUNT(*) FROM claims_claim_lines l "
"LEFT JOIN claims_claims c ON c.claim_id = l.claim_id WHERE c.claim_id IS NULL"
).fetchone()[0]
real = db.execute("SELECT COUNT(*) FROM claims_claims WHERE _dn_provenance <> 'SYNTHETIC'").fetchone()[0]
print(f"claims without a member: {orphans}; lines without a claim: {lines}; non-synthetic rows: {real}")
if orphans or lines or real:
raise SystemExit(1)
print("fixtures ready")Then call it from a workflow:
name: tests-with-synthetic-fixtures
on: [pull_request]
permissions:
contents: read
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- name: Fetch and verify synthetic fixtures
run: python ci/fetch_fixtures.py
- name: Run the tests
run: python -m pytest -qExpected output
verified and loaded enrollment.members: 5 rows
verified and loaded enrollment.coverages: 5 rows
verified and loaded claims.claims: 11 rows
verified and loaded claims.claim_lines: 20 rows
claims without a member: 0; lines without a claim: 0; non-synthetic rows: 0
fixtures readyChange one byte of a downloaded file and the job stops at step 1 with checksum mismatch. In a real suite you would write the database to a file (sqlite3.connect("fixtures.db")) or load the same CSV files into the database service your tests use.
Schema
| Fixture table | Rows | Joined on |
|---|---|---|
enrollment_members | 5 | member_id |
enrollment_coverages | 5 | member_id |
claims_claims | 11 | member_ref → member_id |
claims_claim_lines | 20 | claim_id → claims_claims.claim_id |
SQLite stores the CSV values as text; cast in queries (or declare typed columns from the bundle's schema/*.sql) when a test compares numbers or dates.
What DataNivra does with your own data
Synthetic fixtures cover unit and integration tests. When a pipeline needs data shaped by your own systems, DataNivra's CI integration requests a certified dataset version into an ephemeral environment for the duration of the job and releases it afterwards: the datanivra_ci.py helper and the example workflows in the repository read only CI identifiers from the job environment, and the job's self-hosted runner sits next to your agent so rows never pass through GitHub or DataNivra Cloud. The CI/CD use case describes that flow and its limits. Certification here means DataNivra-certified against configured policy gates; it is not a regulatory or third-party certification and does not make a system compliant with any regulation.
Limits to plan around
- The manifest proves the files are the ones published on the website; it is not a signature. For signed artefacts, verify releases with their cosign bundles as described in the agent release notes.
- Downloading fixtures at build time needs outbound network access from the runner. Cache them (keyed by the manifest digest) if your runners are rate-limited or offline.
- The public bundle is small; it suits integration tests, not performance tests.
Next steps
Read the TDM in CI/CD tutorial, score your pipeline with the CI/CD test-data checklist, and compare with the bundle's own ci/github-actions.yml, which runs tests against DataNivra-certified data built from your own sources.
Synthetic downloads
Files from the Healthcare & Life Sciences pack 2.1.0 bundle. Everything in it is synthetic, generated from a fixed seed, and listed with its SHA-256 digest in the bundle's MANIFEST.json.
synthetic/enrollment.members.csv synthetic/enrollment.coverages.csv synthetic/claims.claims.csv synthetic/claims.claim_lines.csv ci/github-actions.yml
Industry pack: Healthcare & Life Sciences — its entities, scenarios, policy templates and the complete asset bundle.
Try it with DataNivra
The synthetic playground walks through discovery, subsetting, masking, validation and certification in your browser, with no account. Starting free gives you the hosted synthetic sandbox; your own sources need an agent in your network.