Tutorial
Modern delivery pipelines build and deploy many times a day, often into short-lived environments. Test data has traditionally been the part that did not keep up: a shared database refreshed monthly, prepared by hand, and fought over by every branch.
Bringing test data management into CI/CD means a pipeline can ask for the data it needs, get a certified dataset, run its tests and clean up — automatically.
Text description
- CI pipeline (DataNivra Cloud): Commit → Build → Request dataset (CLI/API) → Run tests → Tear down
Connection: Dataset request and status: metadata only
- Your environment: Agent leases the job → Build or reuse certified dataset → Provision ephemeral environment
Why it matters
Shared, long-lived test databases cause flaky tests. One pipeline's test modifies data another pipeline relies on, and failures appear and disappear with no code change. Hand-prepared data also blocks automation: the pipeline is fast until it waits for someone to reload a database. On-demand, isolated datasets remove both problems. They also improve governance, because every request goes through the same policy and certification path as a manual one, instead of pipelines being given broad database access "to make things work".
Example
A pipeline for a synthetic ordering service (all names invented) runs these steps on each merge request:
- Build and unit-test the service.
- Request the dataset product "orders-regression" for an ephemeral environment, with an idempotency key tied to the pipeline run.
- Poll the request until the dataset version is certified and provisioned.
- Run integration tests against the ephemeral environment.
- Tear down the environment; the dataset expires with it.
If certification fails, step 3 reports a failed state and the pipeline stops with a clear reason code, rather than testing against uncertified data. If the pipeline is retried, the idempotency key returns the original request instead of creating a second one.
How DataNivra approaches it
Pipelines use the REST API or the CLI, both of which are metadata-only: they send dataset requests and read states, ids and evidence references. They never connect to source databases and cannot move rows. The request is approved against policy, and the control plane queues a declarative command.
Inside your environment, the agent takes a Lease on the job, builds or reuses a certified version, and handles Provisioning to the ephemeral target. Certification is enforced exactly as for manual requests, and failures block provisioning. Everything row-level remains customer-resident processing, so adding automation does not widen the path for data to leave. See the CLI and API overview for idempotency and error codes.
Checklist for pipeline-ready test data
- Reference datasets by version, not by environment. A pipeline that says "use certified dataset version 12" is reproducible; one that says "use whatever is in QA" is not.
- Fail the build when certification fails. A dataset that did not pass its gates should stop the pipeline, never fall back silently to an older copy.
- Keep secrets out of pipeline logs. Pipelines should pass secret references, never credentials, and print dataset ids and checksums rather than contents.
- Clean up. Deprovision per-run datasets when the run ends so storage and exposure do not grow with every build.
- Separate who can request from who can approve. A pipeline identity should be able to request datasets that policy pre-approves, but changing the masking policy itself should still need a human reviewer, so automation can never widen what it is allowed to see.
- Record which version each test run used. Store the dataset id, version and checksum with the test report, so a failure can be reproduced against exactly the same data weeks later.
- Measure time to data. If requesting a dataset is the slowest step in the pipeline, subset smaller or share one certified version across parallel jobs.
The CI/CD readiness checklist among the free tools turns these points into a self-assessment. The documentation describes the CLI and API used from pipelines, the integrations page lists supported sources, and the guides cover specific databases. Start free to try the workflow against the synthetic sandbox.
Key takeaways
- Give each pipeline run its own certified dataset.
- Use idempotency keys so retries are safe.
- Automation should follow the same policy path as people.