Use cases

Reproducing production issues with masked data

Recreate the records behind a production incident as a small, masked and certified dataset - seed lists stay in your network. A Preview that needs a newer agent release.

Use case

Some defects only appear with the data that triggered them: a customer with an unusual combination of products, a batch of invoices from one particular afternoon, an account migrated from an old system. Engineers then face an uncomfortable choice between guessing from logs and asking for direct production access. DataNivra offers a controlled third path: a reproduction, an issue-recreation request that builds a small, masked, certified dataset around the records involved in an incident. It is a Preview: implemented and tested end to end, and usable once your agent runs a release that supports it.

How a reproduction works

A reproduction is an ordinary dataset request with one extra block: the ticket or incident key from your own tracker (for example OPS-1234), the time window in which the issue occurred (at most 31 days), the bounds of the subset and a short justification. It then goes through the same path as any other dataset: subset, mask, validate, certify, provision.

  • The affected records stay in your network. Your team writes the affected keys into a seed list, a file in a directory your agent is configured to read. The request names it only through an opaque handle such as seedlist://ops-1234; paths and bucket addresses are refused, because they can themselves reveal identifiers. The agent resolves the handle locally, and a key that matches no record fails the build with a count, never the value.
  • Bounded relationship closure. Every reproduction must cap its root records (at most 10,000), its total rows and its fan-out. The subset follows foreign keys and declared cross-system relationships within those bounds, so the related orders, payments or claims come along; if the closure would select more than allowed, the build fails closed instead of quietly growing.
  • No record values in free text. The justification, title and business domain are scanned before anything is stored; an account number or e-mail address typed into them is refused, and filters with literal values are not allowed.
  • Real data only, masked. The dataset must be masked or hybrid. Text columns without a masking rule are redacted by a stricter reproduction profile rather than passed through.
  • Always a second person, always short-lived. Every reproduction waits for approval by someone other than the requester, and its retention is at most 30 days.

Masking and certification are required exactly as for any request that reads real data. Certification is against configured policy gates (DataNivra-certified, not a regulatory certification); the signed certificate records the ticket key, and a reproduction that fails a gate or over-selects is never provisioned.

Where you can use it

  • Console. The New dataset request wizard has a "reproduces a production issue" option with its own issue-recreation step (ticket, window, bounds, justification); the request page shows the reproduction bundle.
  • CLI and SDK. datanivra repro create, repro bundle, repro lease and repro revoke, and the same request through the API and the generated SDKs.
  • Reproduction bundle. A value-free record of the ticket, approvals, seed handle, bounds, policy and graph versions, row counts, certificate and lease, sealed with a SHA-256 digest that the CLI checks. Every read is audited.
  • Hand-off to an ephemeral environment. A reproduction can be leased straight into a short-lived test environment; the lease takes the ticket as its purpose and can never outlive the dataset's retention.
  • Scoped access and revoke by ticket. Only the requester, the approver or a data owner can put a reproduction's rows into an environment, and every grant and refusal is audited. Revoking by ticket cancels pending reproductions, withdraws built ones and ends their leases.

These behaviours are covered by control-plane, agent and engine tests and by an end-to-end test on the real control plane, agent and engine that plants canary values in the seed list and in free text and confirms that none of them reaches any control-plane table, log, bundle or audit record.

What it does not do

  • It needs a newer agent. Agents must advertise the reproduction capability. Published agent releases do not yet include it, so reproductions sent to them are refused before anything is dispatched. It becomes usable when an agent release that includes it is published.
  • Seed lists must be files on the agent host. Lists kept in an object store are not read directly; copy them into the seed-list directory with your own tooling.
  • Closing the incident does not revoke the data automatically. Revoke by ticket from the console, CLI or API, or connect your tracker's close event yourself; reproductions also expire within their retention.
  • It does not restore a source to its state at the time of the incident. The subset reads the source as it is when the job runs, filtered to the window and the seed list.
  • It does not read logs, traces or tickets, and it does not decide which records are affected; your team supplies the seed list or the window.
  • "Small" means the smallest relationship-closed set within your bounds, not an automatically minimised test case.
  • It does not grant production access or bypass masking.

Next steps

Read how subsetting and referential integrity keep related records together, see the ephemeral test environments use case for the lease side, and start free to explore requests against the synthetic sandbox.