Integrations · Files · Available now

Parquet files test data, inside your network

Built-in file-set connector for Parquet (or CSV) files on a local or mounted filesystem; each declared file or directory is a table.

Available now Files

Shipped in the standard agent and backed by recorded conformance evidence.

What works today

Built-in file-set connector for Parquet (or CSV) files on a local or mounted filesystem; each declared file or directory is a table.

Table names come from operator-declared mappings or opaque hashed names, never from data-derived file names; CSV header handling must be declared.

Authentication: filesystem permissions.

Capability matrix

Each capability is tracked separately and shown at the level the recorded conformance evidence supports. A connector that only discovers metadata is never presented as equivalent to one that runs the whole workflow.

Capabilities of Parquet files
CapabilityLevel
DiscoverySupported
ClassificationSupported
ProfilingSupported
MaskingSupported
Deterministic maskingSupported
Entity-aware subsettingSupported
Relationship preservationSupported
Synthetic workflowsSupported
CertificationSupported
ProvisioningSupported
CI/CDSupported
Evidence generationSupported

Read-only proof

Verified from filesystem permissions: the agent user must not be able to write the root.

Conformance evidence

DataNivra runs every connector through a common conformance kit on synthetic data: no write statements, proven read-only access, secret redaction, schema and relationship discovery, bounded streaming, cancellation, clear reason codes, egress canaries and no raw values in logs, then the discover, mask, subset, certify and provision workflow. The published status is computed from those results.

Environments whose conformance run met the preview minimum
Engine or serviceVersionEnvironmentDriverAuthenticationVerified on
Local filesystem (Parquet via Apache Arrow)pyarrow 25.0.1real-enginepyarrow.dataset 25.0.1filesystem permissions2026-09-30

Connector version 0.1.0, last verified 2026-09-30.

What crosses the boundary

Connectors run inside the agent in your network. Only metadata, aggregates and evidence reach DataNivra Cloud:

  • Schema, table and column names and data types (for file sets, operator-declared or opaque table names — never data-derived file names).
  • Aggregates such as row-count estimates, null and distinct ratios; counts from 1 to 10 are reported as 10 (small-cell protection).
  • Classification findings (which column looks like which kind of sensitive data) and relationship metadata.
  • Evidence metadata: checksums, gate outcomes, engine and policy versions, and templated error codes.

Never sent:

  • Any cell value, row, sample, minimum/maximum or top-k value — masked or not.
  • Credentials, connection strings or resolved secret values: DataNivra stores only the secret reference.
  • Free text copied from your data, or exception messages that contain values.

See the customer-resident architecture and how to verify these claims yourself.

Setup outline

  1. Place Parquet files in a directory the agent can read, and mount it read-only into the agent container.
  2. Declare the tables you want (logical name → relative path). Undeclared children get opaque names derived with a keyed hash, so file names that contain data never leave the agent.
  3. Store the connection document in your secret store.
  4. In the console, register the source with the secret reference of that document (for example vault://tdm/source-connection) and the agent that can reach it, then run discovery.

Example connection document, stored in your secret store and resolved by the agent locally (DataNivra only ever sees its reference):

{
  "kind": "PARQUET",
  "root": "/data/extracts",
  "format": "parquet",
  "tables": { "policies": "policies/", "claims": "claims/" }
}

Next: install the agent, then follow the getting started guide.

All integrations · File-based test data management · Guides · Start free