Research
How we benchmark test data pipelines
Benchmarks of data tools are easy to make and hard to trust. This page describes a method anyone can re-run: the estate, what is timed, which checks a run must pass to count at all, and when a number may be quoted.
A deterministic, synthetic estate
The harness generates its own estate from a profile and a seed, so the same command produces the same data on any machine: a relational customer source and a lake-style billing source, five tables and four relationships, one of which crosses the two sources. Three profiles scale the estate from a laptop-sized smoke run to a server-class run of roughly a million rows. Every value is synthetic, and one record carries canary values so that a leak would be detectable. A fingerprint over every generated value lets two people confirm they ran against identical data.
What is measured
- Discovery
- Metadata ingestion, profiling and classification of every column, in wall and CPU seconds.
- Subsetting
- Entity-graph closure across a relational source and a lake-style source joined by a declared cross-source relationship.
- Masking
- An approved keyed policy, with linkage-preserving masking of keys so foreign keys still join after masking.
- Certification
- Validation and every configured certification gate, timed separately from local publishing.
- Referential integrity
- An independent orphan count per relationship on the masked output. It must be zero for a run to count.
- Sensitive-field coverage
- Columns found by classification and masked by policy, compared against the generator’s ground truth.
- Egress
- Every message the agent would send is sealed by the real egress guard with planted canary values; violations and canary hits must be zero.
- Resources
- CPU time, peak memory and storage footprint, recorded together with the machine they were measured on.
When a run counts
A timing is meaningless if the output is wrong. A run is valid only when the dataset was certified and published, every certification gate passed, no relationship lost a row and no canary or egress violation was found; otherwise the harness exits with an error instead of producing a number. Each report records the full environment — profile, seed, source commit and whether the tree was clean, hardware, operating system, virtualisation, Python and package versions — but never a host or user name.
The same approach is used for connector benchmarks on the conformance kit’s synthetic estate and for the cross-system identity graph, each with its own report schema and integrity checks.
Why there are no numbers on this page
Our publishing rule allows a benchmark number outside the repository only together with the complete environment record of a reproducible run from a clean tree on a reference machine. The runs recorded so far come from a developer laptop and are labelled as engineering baselines, so we do not quote them here. A number from one machine is never a sizing promise or a service-level commitment, and we do not compare our numbers with other products’ numbers.
The methodology, harness and example reports are maintained with the product source under docs/benchmarks/; one command (make benchmark) runs a given profile and seed.
Related: the trust and verification page lists the checks you can run yourself, the platform integrity tutorial explains the integrity gates, and the capacity planner estimates storage and refresh costs for your own estate.