Integrations · Lakehouses

Lakehouse test data management

Lakehouses mix open table formats, a catalog and SQL compute. They increasingly hold operational history — customer 360 tables, claims and payment events — alongside analytics.

Two lakehouse features change the test-data picture: table history keeps earlier versions of rows, and the catalog governs access across many workspaces.

What test data management needs here

  • Catalog-aware discovery of catalog.schema.table names, types, constraints and tags.
  • Reads of the current table version only, so historical versions of sensitive rows are never pulled into test data.
  • Service-principal authentication with read-only grants that can be checked before any read.
  • Masking that stays consistent with the operational systems the lakehouse ingests.

How DataNivra approaches it

  • The lakehouse connector design reads through a SQL warehouse as Arrow batches inside your network, with bounded queries.
  • Certification checks masking coverage and relationship integrity across the whole requested dataset, not one table at a time.
  • Whatever the system, row-level work happens in the DataNivra agent inside your network: only metadata, aggregates and evidence reach DataNivra Cloud.

Lakehouses: connector status (1)

1 of these connectors can be used today (available or in preview); the others are shown with their honest status. Statuses are derived from recorded conformance evidence.

Available now

Databricks

Shipped in the standard agent and backed by recorded conformance evidence.

Databricks connector: Unity Catalog catalog.schema.table discovery of Delta tables and views through a Databricks SQL warehouse, informational keys and bounded Arrow reads.

Available now — details about Databricks

Learn, try, then start

See the whole workflow — discovery, classification, masking, subsetting and certification — on synthetic data in the interactive demo, then start a free trial. The product overview explains how the customer-resident agent and DataNivra Cloud divide the work.

Other categories: Databases · Data warehouses · Object storage · Files · NoSQL · SaaS applications · Mainframe · Streaming · APIs.