Integrations · Object storage · Available now

Amazon S3 / S3-compatible storage test data, inside your network

Built-in file-set connector for Parquet or CSV objects under a bucket prefix in Amazon S3 or an S3-compatible store (endpoint override).

Available now Object storage

Shipped in the standard agent and backed by recorded conformance evidence.

What works today

Built-in file-set connector for Parquet or CSV objects under a bucket prefix in Amazon S3 or an S3-compatible store (endpoint override).

Use an IAM policy that grants only s3:GetObject and s3:ListBucket on the prefix.

Authentication: access key (secret reference); instance/workload identity.

Known limitations

  • Verified against Amazon S3 (a real bucket, least-privilege reader) and an S3-compatible server (SeaweedFS S3 gateway); other S3-compatible stores are not individually tested.
  • Only access-key authentication is verified; instance and workload identity are not yet part of the conformance run.
  • Read-only access cannot be proven for object stores: it is attested by the operator, and the source is refused without that attestation.

Capability matrix

Each capability is tracked separately and shown at the level the recorded conformance evidence supports. A connector that only discovers metadata is never presented as equivalent to one that runs the whole workflow.

Capabilities of Amazon S3 / S3-compatible storage
CapabilityLevel
DiscoverySupported
ClassificationSupported
ProfilingSupported
MaskingSupported
Deterministic maskingSupported
Entity-aware subsettingSupported
Relationship preservationSupported
Synthetic workflowsSupported
CertificationSupported
ProvisioningSupported
CI/CDSupported
Evidence generationSupported

Read-only proof

Object-store credentials cannot be introspected portably, so read-only access is by operator attestation (read_only_attested); without it the check fails closed.

Compatibility profiles

Managed variants of an engine are profiles of this connector, not separate connectors. A profile is marked tested only when its own conformance run passed.

  • Amazon S3 — tested

Conformance evidence

DataNivra runs every connector through a common conformance kit on synthetic data: no write statements, proven read-only access, secret redaction, schema and relationship discovery, bounded streaming, cancellation, clear reason codes, egress canaries and no raw values in logs, then the discover, mask, subset, certify and provision workflow. The published status is computed from those results.

Environments whose conformance run met the preview minimum
Engine or serviceVersionEnvironmentDriverAuthenticationVerified on
SeaweedFS S3 gateway (S3-compatible)4.48real-enginepyarrow.fs.S3FileSystem 25.0.1access key (secret reference)2026-09-30
Amazon S3 (amazon-s3-service)servicemanaged-servicepyarrow.fs.S3FileSystem 25.0.1access key (secret reference)2026-10-07

Connector version 0.1.0, last verified 2026-09-30.

What crosses the boundary

Connectors run inside the agent in your network. Only metadata, aggregates and evidence reach DataNivra Cloud:

  • Schema, table and column names and data types (for file sets, operator-declared or opaque table names — never data-derived file names).
  • Aggregates such as row-count estimates, null and distinct ratios; counts from 1 to 10 are reported as 10 (small-cell protection).
  • Classification findings (which column looks like which kind of sensitive data) and relationship metadata.
  • Evidence metadata: checksums, gate outcomes, engine and policy versions, and templated error codes.

Never sent:

  • Any cell value, row, sample, minimum/maximum or top-k value — masked or not.
  • Credentials, connection strings or resolved secret values: DataNivra stores only the secret reference.
  • Free text copied from your data, or exception messages that contain values.

See the customer-resident architecture and how to verify these claims yourself.

Setup outline

  1. Create an identity with s3:GetObject and s3:ListBucket on the bucket prefix only.
  2. Object stores cannot report their permissions portably, so read-only access is by your attestation: set read_only_attested once the policy is in place. Without it, reads that require a read-only source are refused.
  3. Store the connection document with the keys as secret references, or omit the keys so the agent uses the default AWS credential chain of its host (for example an instance or pod role).
  4. In the console, register the source with the secret reference of that document (for example vault://tdm/source-connection) and the agent that can reach it, then run discovery.

Example connection document, stored in your secret store and resolved by the agent locally (DataNivra only ever sees its reference):

{
  "kind": "S3_COMPATIBLE",
  "bucket": "tdm-extracts-example",
  "prefix": "exports/2026-09/",
  "region": "us-east-1",
  "format": "parquet",
  "access_key_id_ref": "vault://tdm/s3-reader#access_key_id",
  "secret_access_key_ref": "vault://tdm/s3-reader#secret_access_key",
  "read_only_attested": true
}

Next: install the agent, then follow the getting started guide.

All integrations · Object storage test data management · Guides · Start free