Lakehouse test data · Available now
Test data management for Databricks, inside your network
Lakehouses now hold operational data — customer 360 tables, claims histories, payment events — and data engineers need realistic data to build and test pipelines. DataNivra brings its customer-resident masking, subsetting and certification workflow to Databricks without copying production rows out of your environment.
Available now. Shipped in the standard agent and backed by recorded conformance evidence. The Databricks connector is available now: it ships in the standard agent and passed every conformance check.
Using DataNivra with Databricks
The Databricks connector is available now: it ships in the standard agent and passed every conformance check.
- Unity Catalog discovery Supported
- Discovery walks the three-level catalog.schema.table namespace that the service principal can browse, and records column names, types, informational primary and foreign key constraints, table comments and tags as metadata. Tags such as a sensitivity label are useful classification hints; they are read, never written back.
- Delta tables through a SQL warehouse Supported
- Reads are ordinary SELECT statements executed on the Databricks SQL warehouse you configure, fetched as Arrow batches by the agent inside your network. Delta time travel is not used to read history: the connector reads the current version of each table, so older versions of sensitive rows are never pulled into a test dataset.
- OAuth service principal and secret references Available now
- The design authenticates as a Databricks service principal with OAuth machine-to-machine credentials (or a personal access token), both supplied as secret references that the agent resolves locally. DataNivra Cloud stores only the reference, never the secret.
- Proven read-only access Available now
- Before reading, the connector checks the grants of its principal: only USE CATALOG, USE SCHEMA, SELECT and BROWSE are acceptable. MODIFY, CREATE or ALL PRIVILEGES stops the job, so a mis-scoped principal fails closed instead of reading with write rights.
- Bounded, cost-conscious reads Supported
- Discovery reads metadata only. Profiles and subset reads run queries on your warehouse, bounded by row limits and a statement timeout, so a test-data job cannot turn into an unbounded scan of a production lakehouse.
- Deterministic masking across systems Supported
- Masking runs in the agent with keys that stay in your environment. Deterministic, keyed pseudonymisation gives the same customer the same masked key in Databricks and in the operational databases it came from, so joins between a lakehouse table and a PostgreSQL or SQL Server table still work in test.
- Entity-aware subsets Supported
- Unity Catalog does not enforce foreign keys, so relationships come from informational constraints where declared and from explicit relationship mappings you approve. The subsetter then follows them to pull a customer with the claims, policies and events that belong to them, and nothing unrelated.
- Certification before provisioning Supported
- A masking job that finished is not a certified dataset. Certification checks masking coverage, relationship integrity, policy versions, row counts and the manifest, and provisioning refuses any dataset that failed or was revoked.
- CI/CD and evidence Supported
- Pipelines request certified datasets through the REST API, and each dataset carries an evidence record — policy and engine versions, gate results and checksums, never row values — that an auditor can review later.
Not in scope for the first version: external locations, volumes, Delta Sharing and Databricks Jobs. They will be added only after they are tested.
Why production lakehouse copies are risky
Development and production workspaces often reach the same storage accounts, so a development notebook reading a production table is one grant away. Delta time travel keeps earlier versions of every row until the retention period passes and files are vacuumed, so masking a table in place leaves the original values reachable. Wide, denormalised tables mix identifiers, free text and derived features across hundreds of columns. A test-data process for Databricks therefore has to classify column by column, read only the current version, and deliver a separate, certified dataset rather than an edited copy.
Read the longer Databricks test data guide and the lakehouse overview.
Get started
Create a service principal with read-only grants, store its OAuth credentials in your secret store, and register the source with its secret reference. The Databricks integration page has the capability matrix, tested environments and known limitations.
Capabilities are tracked one by one: Discovery, Classification, Profiling, Masking, Deterministic masking, Entity-aware subsetting, Relationship preservation, Synthetic workflows, Certification, Provisioning, CI/CD, Evidence generation. Each is shown at the level its conformance evidence supports.