Guide
Azure Data Lake Storage Gen2 is where many Azure estates keep raw and curated data: exports from operational databases, files from partners, and the bronze layer of analytics platforms. It is a practical source of realistic test data, as long as sensitive files are not copied into development subscriptions along the way. This guide shows how to use DataNivra's built-in Azure Blob / Data Lake connector.
What the connector supports
The Azure Blob / Data Lake connector ships with the DataNivra agent. It reads Parquet or CSV files under a prefix in an Azure Blob container or an ADLS Gen2 filesystem; each declared file or directory is a table. It authenticates with an account key, a SAS token or a service principal (tenant id, client id and client secret), all supplied as secret references that the agent resolves locally. Custom blob and Data Lake endpoints are supported, which also makes local testing against an emulator possible. See the Azure Blob / Data Lake integration page.
Read-only access is attested, not verified
Storage permissions cannot be introspected portably, so the agent cannot prove that a credential is read-only the way it can for PostgreSQL. Instead, you attest it with read_only_attested after scoping the identity; without the attestation the read-only check fails closed and production-flagged jobs do not read. Make the attestation true:
- Prefer a service principal with only the Storage Blob Data Reader role, scoped to the container rather than the whole account.
- Avoid account keys where you can: a key grants full control of the account.
- If you use SAS, issue read and list permissions only, with a short expiry, and rotate it through your secret store.
- Enable storage diagnostic logs so you can later show that only reads occurred.
Keeping the data inside your network
Deploy the agent in your own Azure subscription, in a VNet with a private endpoint or service endpoint for the storage account, so data never crosses the public internet. The agent makes outbound HTTPS calls to DataNivra Cloud and accepts no inbound connections. The repository includes a customer-VNet package for Azure, and the documentation walks through installation and secret providers such as Key Vault.
What leaves the VNet is metadata only: table and column names, data types, row counts, ratios, gate outcomes and checksums. Rows, samples and credentials never do; the zero raw-production-data egress tutorial explains how that is enforced in the agent and again at the control plane.
File names, headers and relationships
The same rules apply as for every file source. Data-derived file and directory names are never reported: declare logical table names, or let the agent expose opaque hashed names. CSV sources must declare header handling; the agent will not guess. And because Parquet and CSV carry no constraints, declare relationships through policy or an industry pack so the subset engine can keep them intact.
A typical workflow
- Land exports from operational systems into a dedicated container prefix, with lifecycle rules that delete them after the dataset is certified.
- Register the prefix as a source, attest read-only access, and run discovery.
- Review classifications, then build an entity-based subset. The subset-size calculator among the free tools helps estimate volume.
- Mask with deterministic, keyed strategies so the same customer receives the same pseudonym everywhere, then certify. Datasets that fail any gate are never provisioned; see dataset certification.
- Provision the certified version to a directory or a PostgreSQL test database for your QA environment.
An interim path for other sources
Many Azure estates also run SQL Server or Databricks. SQL Server has its own connector in preview, with the limitations listed in the SQL Server guide (Azure SQL Database and Managed Instance are not yet tested). The Databricks connector is not shipped yet, so exporting its tables as Parquet into ADLS and reading them here remains the interim path; see the Databricks guide. Every connector's status is on the integrations page.
Next steps
Start free with the synthetic sandbox, then connect a test container. The customer-resident processing tutorial is a good read for the security review that usually precedes connecting a production data lake.