Data subsetting
How to carve small, representative, referentially complete slices out of large databases, driven by the business entities your tests care about.
Topics
A subset is useful only if it is consistent: every order needs its customer, every claim its policy, across databases and files. These pages explain how entity-based subsetting keeps those relationships intact, how to size a subset before building it, and how a narrow subset helps investigate a production incident.
How to carve small, representative, referentially complete slices out of large databases, driven by the business entities your tests care about.
Why every foreign key in a test dataset must resolve, and how masking and subsetting can quietly break joins across tables and systems.
How to give QA one customer and every related policy, claim, account and transaction across several systems, without dangling references or unrelated records.
Recreate the records behind a production incident as a small, masked and certified dataset - seed lists stay in your network. A Preview that needs a newer agent release.
A runnable comparison of naive row sampling and an entity-driven subset on synthetic claims data, counting the dangling references each one leaves behind.