Use cases

Unstructured and AI test data

Mask documents, chat and ticket exports inside your network and generate synthetic RAG evaluation sets - a Preview with rule-based detection and its limits stated.

Use case

Tables are only part of what a modern application stores. Support tickets, chat transcripts, e-mail threads and case notes hold the same names, phone numbers and account numbers as the customer table, and teams building search, assistants or retrieval-augmented generation (RAG) systems need realistic documents to test with. Copying those documents into a test environment, or sending them to a model provider, moves personal data out of the place it is governed. DataNivra's unstructured test data capability is a Preview: it is implemented and tested end to end with synthetic documents, and it has known limits that this page states plainly.

How document sets are masked

A document set is a directory or object-store prefix of documents that your agent reads inside your network. Plain text, Markdown and log files, JSON, JSON Lines chat and ticket exports, e-mail messages and Word (DOCX) files are parsed by hardened readers that reject oversized, malformed or macro-enabled input. The set appears to the build as one table, documents, with one row per document part.

A masking policy with a TEXT_SPAN_MASK rule finds personal and sensitive spans inside the text - names, e-mail addresses, phone numbers, national, payment and bank identifiers, medical record and provider numbers, dates of birth, street addresses and cities - and replaces each one in place. Every entity type uses an ordinary keyed masking strategy with the rule's own key reference, so the same person gets the same surrogate in a ticket and in the customer table masked with the same key. After replacement, a residual re-scan runs the detectors again over the masked text; anything detectable outside the inserted surrogates fails the build and nothing is published.

The dataset then goes through the same path as any other request: approval, validation, certification against configured policy gates (DataNivra-certified, not a regulatory certification) and provisioning. A document-set version gets no certificate without a passing document-set summary.

What reaches DataNivra Cloud

Only counts, codes, versions and SHA-256 digests: how many documents and parts were read, rejections per reason code, detected spans per entity type with small counts bucketed, the residual re-scan outcome, the detector-set version and the model-provider mode in force. Document text, spans, offsets, file names and paths stay in your environment and in your evidence bundle.

Synthetic RAG evaluation sets

An evaluation set tests a retrieval-augmented system: each case has a question, the documents that should ground the answer, a golden answer, grading criteria, the asker's authorization context and adversarial tags for prompt injection, planted canaries, access boundaries and unanswerable questions. Sets are generated customer-side from a seed and templates, never from your documents, and the same seed and versions produce the same bundle. DataNivra Cloud registers only the manifest, so a team can check later that the bundle it tests with is the one that was registered.

Model providers stay off by default

No model receives customer content by default; the detectors run in the agent process and call nothing. A tenant administrator who is neither the author nor the submitter can approve a model-provider policy that names exact endpoints, purposes and an expiry, and the agent re-checks it before any call. No model adapter ships today, so even an approved policy still refuses every model call.

What it does not do

  • It does not prove anonymization. Detection is deterministic and rule-based: patterns, checksums, keyword context, a name lexicon, a city list and a per-document consistency pass. It still misses values it has no rule or context for, such as a bare first name in narrative text or a city in unusual phrasing, and it can over-mask product names, order numbers and dates. The residual re-scan uses the same detectors, so it cannot prove that nothing was missed. Expect lower recall on real text than on our synthetic test corpora, and review samples of masked output in your own environment.
  • No published accuracy figures. DataNivra measures the detector on its own synthetic corpora for engineering purposes and does not publish those numbers as a product claim.
  • No PDF or legacy Office files. They are rejected; a PDF extractor is not part of the Preview.
  • No re-rendered files. Masked documents are delivered as the documents table, one row per part, not as rewritten originals.
  • No model calls. No model adapter ships, and choosing a provider and its data processing terms is a decision for your organization.
  • A newer agent is required. Requests need an agent release that advertises the document-set and in-text masking capabilities; older agents are refused before anything is sent.

Next steps

Read the developer reference for unstructured test data for the formats, rule parameters, evidence fields and commands, the data masking tutorial for how keyed masking keeps identities consistent, and pseudonymization versus anonymization for why masked text is not anonymous data.