Integrations · Files · Available now
Parquet files test data, inside your network
Built-in file-set connector for Parquet (or CSV) files on a local or mounted filesystem; each declared file or directory is a table.
Available now Files
Shipped in the standard agent and backed by recorded conformance evidence.
What works today
Built-in file-set connector for Parquet (or CSV) files on a local or mounted filesystem; each declared file or directory is a table.
Table names come from operator-declared mappings or opaque hashed names, never from data-derived file names; CSV header handling must be declared.
Authentication: filesystem permissions.
Capability matrix
Each capability is tracked separately and shown at the level the recorded conformance evidence supports. A connector that only discovers metadata is never presented as equivalent to one that runs the whole workflow.
| Capability | Level |
|---|---|
| Discovery | Supported |
| Classification | Supported |
| Profiling | Supported |
| Masking | Supported |
| Deterministic masking | Supported |
| Entity-aware subsetting | Supported |
| Relationship preservation | Supported |
| Synthetic workflows | Supported |
| Certification | Supported |
| Provisioning | Supported |
| CI/CD | Supported |
| Evidence generation | Supported |
Read-only proof
Verified from filesystem permissions: the agent user must not be able to write the root.
Conformance evidence
DataNivra runs every connector through a common conformance kit on synthetic data: no write statements, proven read-only access, secret redaction, schema and relationship discovery, bounded streaming, cancellation, clear reason codes, egress canaries and no raw values in logs, then the discover, mask, subset, certify and provision workflow. The published status is computed from those results.
| Engine or service | Version | Environment | Driver | Authentication | Verified on |
|---|---|---|---|---|---|
| Local filesystem (Parquet via Apache Arrow) | pyarrow 25.0.1 | real-engine | pyarrow.dataset 25.0.1 | filesystem permissions | 2026-09-30 |
Connector version 0.1.0, last verified 2026-09-30.
What crosses the boundary
Connectors run inside the agent in your network. Only metadata, aggregates and evidence reach DataNivra Cloud:
- Schema, table and column names and data types (for file sets, operator-declared or opaque table names — never data-derived file names).
- Aggregates such as row-count estimates, null and distinct ratios; counts from 1 to 10 are reported as 10 (small-cell protection).
- Classification findings (which column looks like which kind of sensitive data) and relationship metadata.
- Evidence metadata: checksums, gate outcomes, engine and policy versions, and templated error codes.
Never sent:
- Any cell value, row, sample, minimum/maximum or top-k value — masked or not.
- Credentials, connection strings or resolved secret values: DataNivra stores only the secret reference.
- Free text copied from your data, or exception messages that contain values.
See the customer-resident architecture and how to verify these claims yourself.
Setup outline
- Place Parquet files in a directory the agent can read, and mount it read-only into the agent container.
- Declare the tables you want (logical name → relative path). Undeclared children get opaque names derived with a keyed hash, so file names that contain data never leave the agent.
- Store the connection document in your secret store.
- In the console, register the source with the secret reference of that document (for example vault://tdm/source-connection) and the agent that can reach it, then run discovery.
Example connection document, stored in your secret store and resolved by the agent locally (DataNivra only ever sees its reference):
{
"kind": "PARQUET",
"root": "/data/extracts",
"format": "parquet",
"tables": { "policies": "policies/", "claims": "claims/" }
}Next: install the agent, then follow the getting started guide.
All integrations · File-based test data management · Guides · Start free