Guide
Snowflake makes it trivially easy to create a development copy of a warehouse: a clone is instant and costs nothing until the data diverges. That convenience is also the risk. A cloned production database carries every sensitive column into an environment with broader access, and nobody has to approve it. This guide explains how to think about Snowflake test data and where DataNivra's Snowflake connector stands today.
Current status: available
The Snowflake connector is available in agent 0.3.0 and later, whose standard image includes the Snowflake driver. It passed DataNivra's connector conformance kit against a real Snowflake account, covering discovery, the read-only proof, bounded reads, masking, subsetting, certification and provisioning of a synthetic estate, and a role with INSERT was refused.
Known limitations include agent 0.3.0 or later being required (earlier agent images do not include the driver), relationships that Snowflake does not enforce (they may need explicit mappings) and NUMBER(p,0) values beyond 64-bit integers, which fail the job with a reason code. The Snowflake integration page shows the live status from the product's connector registry, and this guide is reviewed whenever it changes.
Clones are not test data
Snowflake's clone feature solves storage cost, not exposure. A clone of production:
- contains the same personal and financial data as the source, unless masking policies are applied, and those policies mask values at query time by role rather than changing the stored data, so a role with broader access still sees the originals;
- inherits no business intent: it is every row, not the entities a test needs;
- is invisible to most test-data governance, because no data was "copied" in the traditional sense.
Masked subsets are different: a smaller set of connected records in which sensitive values have actually been replaced, with evidence that the replacement happened. The synthetic vs masked data tutorial covers when to use masked production-like data and when to generate synthetic data instead.
The read-only proof the connector requires
The rule is strict where it matters. The connector turns off secondary roles for its session. It then reads SHOW GRANTS TO ROLE for the extraction role, for every role that role inherits, and for PUBLIC. On the source database and everything in it, those grants may contain only USAGE, SELECT, REFERENCES, MONITOR and READ. Any INSERT, UPDATE, DELETE, OWNERSHIP or CREATE privilege there stops the job, and so do an administrative role such as SYSADMIN, any privilege on a role or user, and account-level privileges such as MANAGE GRANTS or APPLY MASKING POLICY. Snowflake's own defaults for PUBLIC (feature switches such as VIEW LINEAGE and USE AI FUNCTIONS, and roles in Snowflake's system database) cannot change your data and do not block the source. You can set up that role now: a dedicated role with usage on one small warehouse, the database and the schema, and select on the required tables. Authenticate the user with key-pair authentication.
Alternative: unload to Parquet in your own storage
If you prefer not to connect the agent to Snowflake directly, unload data to storage you control and read it with a shipped object-store connector:
- Use
COPY INTO @stagewithFILE_FORMAT = (TYPE = PARQUET)to unload the tables you need to an external stage in your own Amazon S3 bucket or Azure container. Filter at unload time to the entities in scope rather than exporting whole tables. - Register that location as an Amazon S3 or Azure Blob / Data Lake source. Object-store read-only access cannot be introspected, so the agent requires you to attest that its credential is read-only (for example an IAM policy with only
s3:GetObjectands3:ListBucket); without the attestation the check fails closed. - Declare table names for each unloaded prefix, and declare the relationships between tables, since Parquet carries no constraints.
- Subset, mask, certify and provision. Certified datasets can be written to a directory or a PostgreSQL target today; loading them back into a Snowflake development database is a
COPY INTOyour pipeline runs from the certified files.
The unload files are sensitive until masked: keep the stage private, encrypted and short-lived.
Cost and cadence
Unloading, masking and re-loading costs warehouse credits and storage, so refresh cadence matters more on Snowflake than on a small database. Subset first, then mask: processing a few percent of the rows is cheaper than masking a full clone. The storage and compute tutorial and the refresh cadence tutorial help you pick a rhythm, and the storage calculator among the free tools lets you compare scenarios with your own assumptions.
How DataNivra approaches it
All row-level processing runs in the agent inside your cloud account or network. DataNivra Cloud receives schema and column names, counts and evidence references only. Certification blocks any dataset whose masking coverage or referential integrity cannot be proven.
Next steps
- Track Snowflake and other connectors on the integrations page.
- Read the documentation on object-store sources.
- Start free with the synthetic sandbox to evaluate the workflow before building the unload job.