Guide
Amazon S3 has become the landing zone for almost every data flow: database exports, event archives, partner feeds and the raw layer of data lakes. That makes it a natural source of realistic test data, and a common place for sensitive files to be copied without anyone noticing. This guide explains how to build test datasets from S3 safely with DataNivra's built-in S3 connector.
What the connector supports
The Amazon S3 / S3-compatible connector ships with the DataNivra agent and is available. It reads Parquet or CSV files under a bucket prefix, in Amazon S3 or in an S3-compatible store reached through an endpoint override. Each declared file or directory becomes a table. The agent streams rows as Arrow batches into the local engine for classification, subsetting, masking and synthesis.
It passed every conformance check and the full workflow against a real Amazon S3 bucket, read by an IAM user allowed only s3:GetObject on the data prefix and s3:ListBucket limited to that prefix, and against an S3-compatible object store. Read-only access to object storage is by operator attestation, because credentials cannot be introspected portably: the agent refuses the source until you attest it. The Amazon S3 integration page lists the current limitations and the tested environments.
Credentials are supplied as secret references and resolved by the agent: an access key, secret key and optional session token, or whatever your secret provider holds. They never reach DataNivra Cloud.
Read-only access is attested, not proven
For a database, the agent can prove from the catalog that a credential cannot write. For an object store it cannot: there is no portable way to ask S3 what an access key is allowed to do. The connector is honest about that. Read-only access is established by operator attestation: you set read_only_attested after scoping the credential, and without the attestation the read-only check fails closed with READ_ONLY_UNVERIFIABLE_STORE.
Make the attestation true before you set it:
- Grant only
s3:GetObjecton the prefix ands3:ListBucketon the bucket, limited to that prefix. - Do not grant
s3:PutObject,s3:DeleteObjector any bucket-policy permissions. - Prefer a dedicated IAM role for the agent (for example through the instance or task role in your VPC) over long-lived keys.
- Turn on S3 server access logging or CloudTrail data events for the prefix so you can confirm, later, that only reads happened.
Table names and CSV headers
File names are often data: an export called after a customer or an account number is a leak if it becomes metadata. DataNivra therefore never reports data-derived file names. Either you declare logical table names mapped to relative paths, or each child of the prefix is exposed under an opaque hashed name, with the mapping kept on the agent host.
CSV files must declare whether the first line is a header. With has_header: true the header is accepted only when every name looks like an identifier; with false, columns are named positionally or from names you declare. Discovery refuses an undeclared header rather than guessing, because a headerless first line is a data record.
Relationships across files
Parquet and CSV carry no foreign keys. To build referentially intact subsets from files, declare the relationships through a policy or an industry pack, for example claims.policy_id references policies.policy_id. The engine then follows them exactly as it would database constraints. The referential integrity tutorial explains why this matters.
Deploying the agent next to the bucket
Run the agent inside your AWS account, in a VPC with an S3 gateway endpoint, so reads never cross the public internet. The agent needs outbound HTTPS to DataNivra Cloud and nothing inbound. The repository includes a customer-VPC package for AWS, and the documentation covers installation. Everything row-level stays in that VPC; DataNivra Cloud receives table and column names, counts, gate outcomes and checksums only.
From files to a certified dataset
- Register the prefix as a source and run discovery; review the proposed classifications.
- Choose root entities and build a subset; use the subset-size calculator among the free tools to estimate volume first.
- Apply masking: synthetic replacements for names, emails and addresses; deterministic pseudonyms for keys so joins still work.
- Certify. Gates check policy coverage, masking completion, integrity, row counts and checksums, and a failing dataset is never provisioned. See dataset certification.
- Provision to a directory target or a PostgreSQL test database.
Also useful for other sources
Because many databases can export Parquet to S3, this connector is also the interim path for sources whose connectors are not shipped yet, such as Snowflake and Databricks. The integrations page lists every connector with its status. Ready to try it? Start free with the synthetic sandbox, then connect a test bucket.