Tutorial
Subsetting creates a smaller dataset that still behaves like the whole. Instead of cloning ten terabytes, you select the records your tests actually need and bring along everything those records depend on.
The key idea is to start from a business entity — a member, a customer, an order — rather than from tables. Tests are written about entities, so subsets should be defined that way too.
Text description
- Subsetting (Your environment): Select driving entities → Follow keys to parents → Include required children → Close the reference set → Verify no dangling keys
Why it matters
Smaller datasets restore faster, cost less to store and reduce the volume of sensitive data sitting in test environments. They also make tests more predictable: a curated set of 500 members with known characteristics is easier to reason about than an entire population.
But careless subsetting is worse than none. Take 1% of each table independently and you get claims whose members do not exist and accounts with no customer. Applications crash or, worse, silently skip records. A useful subset must reach referential closure.
Example
A synthetic subset definition for a claims system:
driving entity : member
filter : plan_type = 'PPO' and has_claim_since = today - 90 days
sample : 500 members, stratified by state
include : coverage, claims, claim_lines, providers (referenced)
exclude : audit_log (not needed for these tests)Starting from 500 members, the subsetter follows foreign keys up to required parents (plans, providers referenced by claims) and down to required children (claims, claim lines). If a claim references a provider, that provider is included even if it serves members outside the subset. The result is then checked so that no key dangles. Table names and figures are illustrative.
How DataNivra approaches it
In DataNivra a subset is a declarative definition — driving entity, filters, date windows, sampling and relationship rules — attached to a dataset request and approved like any other policy. The agent executes it inside your environment using the relationships from schema metadata plus the entity model from an industry pack, which fills in relationships that databases do not declare.
The subset is then masked or supplemented with synthetic records, and certification checks referential integrity before provisioning. Because the definition is versioned, the same subset can be rebuilt on refresh. Only counts and evidence references return to the control plane. See Referential integrity next.
Key takeaways
- Subset by business entity, not by table.
- Follow relationships until nothing dangles.
- Keep subset definitions versioned so results are repeatable.