Guides

Test data management for Databricks

Test data for Databricks lakehouses - Delta tables, Unity Catalog grants, masked subsets - and the current status of DataNivra's Databricks connector.

Guide

Databricks lakehouses increasingly hold operational data, not just analytics: customer 360 tables, claims histories, transaction feeds. Data engineers need realistic data to develop and test pipelines, and the easy answer is to point the development workspace at production tables. This guide explains what good Databricks test data looks like and where DataNivra's Databricks connector stands today.

Current status: available

The Databricks connector is available in agent 0.3.0 and later, whose standard image includes the Databricks driver. It reads Unity Catalog catalog.schema.table objects through a Databricks SQL warehouse, and it passed DataNivra's connector conformance kit against a real workspace on a serverless SQL warehouse, covering discovery, the read-only proof, bounded reads, masking, subsetting, certification and provisioning of a synthetic estate; a service principal with MODIFY was refused.

Known limitations: agent 0.3.0 or later is required (earlier agent images do not include the driver), and workspace and metastore administrator rights are invisible to the read-only proof, so use a dedicated, non-admin service principal. The Databricks integration page shows the live status from the product's connector registry.

Why lakehouse test data needs care

  • Shared storage, shared risk. Development and production workspaces often reach the same storage accounts. A development notebook reading a production table is one grant away.
  • Wide tables. Denormalised tables mix identifiers, free text and derived features. Classification has to work column by column across hundreds of columns.
  • Time travel keeps the past. Delta tables retain previous versions. Masking a table in place leaves the original values reachable through history until the retention period expires and the files are vacuumed.
  • Pipelines need edge cases. Testing a pipeline properly needs nulls, late-arriving records and boundary values, which a random sample rarely contains.

The read-only proof the connector requires

The connector authenticates as a service principal with OAuth machine-to-machine credentials, supplied as secret references. It then reads the Unity Catalog privilege views for that principal and every group it belongs to. It accepts only USE CATALOG, USE SCHEMA, SELECT and BROWSE, and the principal must own no catalog, schema or table. MODIFY, CREATE TABLE, ALL PRIVILEGES or any ownership stops the job. Workspace administrator rights are not visible through those views, so use a dedicated, non-admin service principal.

Alternative: write Parquet to storage you control

If you prefer not to connect the agent to a SQL warehouse, write the data you need as Parquet to cloud storage in your own account, then read it with a shipped connector:

  1. In a job with the read-only principal, select the entities in scope and write them as plain Parquet (not Delta) to a dedicated location in your Amazon S3 bucket or Azure Data Lake Storage container. Plain Parquet avoids transaction logs and time travel in the extract.
  2. Register the location as an Amazon S3 or Azure Blob / Data Lake source. Because object-store permissions cannot be introspected portably, the agent requires an explicit read-only attestation for the credential and fails closed without it.
  3. Declare a logical table name for each directory and declare the relationships between tables.
  4. Discover, classify, subset, mask and certify. Certified outputs can be provisioned to a directory target, from which your pipeline loads development tables.

Delete the extract once the certified dataset exists; it is sensitive until masked. Treat the extract location like production: restrict it to the extraction principal and the agent's identity, enable storage logging, and add a lifecycle rule that removes files after a short period so a forgotten extract does not become a second, unmanaged copy of production data.

Masked data or synthetic data?

Pipeline tests often need volume and edge cases more than they need production-like distributions. DataNivra's synthetic generator produces scenario-labelled rows (normal, boundary, rare, null-heavy, high-volume, historical, and clearly labelled negative-test data) from a schema, without reading source values. Masked subsets are better when the logic under test depends on real-world correlations. The synthetic vs masked data tutorial compares the two, and the scenario planner among the free tools helps you choose a mix.

How DataNivra approaches it

The agent runs inside your environment and reads from storage you own. Only metadata such as table and column names, counts and evidence references reach DataNivra Cloud. Each certified dataset has a manifest and checksums, so a pipeline can prove which version it tested against; see dataset certification.

Next steps