Documentation

Unstructured and AI test data (Preview)

Document sets, in-text masking with TEXT_SPAN_MASK, value-free document-set evidence, synthetic RAG evaluation sets and the default-deny model-provider policy — and their limitations.

Status

Unstructured test data is a Preview capability. Everything on this page is implemented and tested end to end with synthetic documents, but detection is deterministic and rule-based and has known blind spots (see Limitations below). Masking free text reduces exposure; it is not proof of anonymization.

Document sets

A document set is a directory (or object-store prefix) of documents that the agent reads inside your network. Supported formats:

FormatWhat is extracted
.txt, .md, .logThe whole text (strict UTF-8).
.jsonEvery string value; JSON keys are never read as text.
.jsonl, .ndjsonChat and ticket exports; the fields named in document_text_fields, or every string.
.emlSender, recipient and subject headers, and the plain-text and HTML parts. Attachments are skipped.
.docxBody, headers, footers, footnotes, endnotes and comments. Macro-enabled files are rejected.

PDF and legacy Office formats are not parsed; they are rejected with DOCUMENT_TYPE_UNSUPPORTED. The parsers are hardened against zip bombs, XML entity expansion, path traversal and oversized input, and a document they cannot read safely is rejected with DOCUMENT_PARSE_REJECTED. By default one rejected document fails the build; document_on_rejected: skip leaves rejected documents out and records how many there were.

To use a document set:

  1. In the agent's connection document (in your secret store) use a LOCAL_FILES, PARQUET, S3_COMPATIBLE or AZURE_BLOB connection with format: documents. Set table_name_key so document ids are opaque.
  2. Register the source with content_format: DOCUMENTS. The control plane only ever knows the source name, kind and secret reference.
  3. Mask it with a policy that has a TEXT_SPAN_MASK rule on the text column of the documents table. A request from a DOCUMENTS source without such a rule is refused with TEXT_SPAN_RULE_MISSING.

The set is exposed to the build as one table, documents, with one row per document part (doc_id, part_index, part_name, media_type, text).

In-text masking: TEXT_SPAN_MASK

TEXT_SPAN_MASK finds personal and sensitive spans inside free text — names, e-mail addresses, phone numbers, national and payment identifiers, IBANs, medical record numbers, provider numbers, dates of birth, street addresses, IP addresses and credentials — and replaces each span in place. Each entity type uses an ordinary masking strategy (for example NAME_REPLACE for names and FORMAT_PRESERVING_SYNTHETIC for identifiers) with the rule's own key reference, so the same person gets the same surrogate in documents and in tables masked with the same key.

Rule parameters:

  • ENTITY_<TYPE>=<strategy> overrides the strategy for one entity type. Strategies that cannot replace a span (PASS_THROUGH, NULLIFY, NUMERIC_PERTURB, CUSTOM_PLUGIN) are refused with TEXT_SPAN_STRATEGY_UNSUPPORTED.
  • MIN_CONFIDENCE (0 to 1) and MAX_TEXT_CHARS (larger values fail with TEXT_TOO_LARGE).

After replacement, a residual re-scan runs every detector again on the masked text. Any detection outside the inserted surrogates fails the build with UNSTRUCTURED_RESIDUAL_PII; nothing is published.

Evidence and certificates

For a document-set build the agent reports a document-set summary with its certification report: status, document and part counts, rejected documents per reason code, detected spans per entity type, the residual re-scan outcome, the detector-set version and SHA-256, the detector adapters used (none by default), the model-provider mode in force and two gates, UNSTRUCTURED_RESIDUAL_PII and DOCUMENT_PARSE. It never contains text, spans, offsets, file names or paths. Counts describe source content, so small counts (1 to 10) are reported as 10. The summary's evidence_sha256 digests the agent's local evidence, which stays in your evidence bundle.

The signed test-data certificate of such a version carries the same information as document_set. A version built from a DOCUMENTS source gets no certificate without a passing summary (DOCUMENT_SET_SUMMARY_MISSING). Certificates issued before this field existed verify exactly as before. Read the summary with GET /v1/datasets/{dataset_id}/versions/{version}/document-set.

Agent versions

Document sets and TEXT_SPAN_MASK need an agent that advertises the DOCUMENT_SETS and TEXT_SPAN_MASK capabilities. Older agents parse commands strictly and would reject them, so the control plane refuses such requests and jobs with AGENT_DOCUMENT_SETS_UNSUPPORTED or AGENT_TEXT_SPAN_MASK_UNSUPPORTED instead of sending them. Upgrade the agent and request again.

Model-provider policy

No model receives customer content by default. The deterministic detectors run in the agent process and call nothing.

  • Each organization has at most one MODEL_PROVIDER policy. Its mode is NONE (the default), LOCAL_ONLY or APPROVED_EXTERNAL.
  • An APPROVED_EXTERNAL version lists exact https endpoints, the purposes they may serve, an expiry and a credential reference. It must be submitted for review and approved by a tenant administrator who is neither its author nor its submitter (TENANT_ADMIN_REQUIRED, separation of duties). Every decision is audited.
  • The approved version travels to the agent inside the signed TEXT_SPAN_MASK command. The agent's model egress gate re-checks mode, approval, separation of duties, expiry, purpose and the endpoint allow-list, and denies everything else (MODEL_EGRESS_*).
  • No model adapter ships. Even under an approved policy, every model call is refused with MODEL_ADAPTER_NOT_AVAILABLE. Choosing a provider, its data processing agreement and sub-processor listing are decisions for your organization and DataNivra's owner.
datanivra model-policy show                      # effective mode; NONE unless approved
datanivra model-policy create --name "Model providers"
datanivra model-policy propose --file policy.json
datanivra model-policy submit 1
datanivra model-policy approve 1 --reason "DPA signed" --expected-checksum <sha256>

Synthetic RAG evaluation sets

An evaluation set tests a retrieval-augmented system: each case has a question, the documents that should ground the answer, a golden answer and grading criteria, adversarial tags (PROMPT_INJECTION, PII_LEAK_CANARY, ACCESS_BOUNDARY, UNANSWERABLE) and the asker's authorization context. Sets are synthetic: generated from seeded templates, never from your documents.

datanivra-agent eval-set generate --seed 2026 --out eval.json   # customer-side, offline
datanivra eval-set verify eval.json                             # digests, exit 32 on mismatch
datanivra eval-set register eval.json --name "support-bot v1"   # registers the manifest only
datanivra eval-set export <eval-set-id> --out manifest.json     # manifest + how to regenerate
datanivra eval-set verify eval.json --eval-set-id <eval-set-id>

The same seed and versions produce the same bundle. The control plane stores only the manifest — versions, seed, case and tag counts and SHA-256 digests — and never case text.

Local scan

datanivra documents scan --dry-run PATH runs the agent's parsers and detectors on a local directory (through datanivra-agent documents scan) and prints counts only: documents, parts, media types, rejections per reason code and spans per entity type. Nothing is uploaded and nothing is written.

Limitations

  • Rule-based detection. Detection uses patterns, checksums, keyword context, a name lexicon, a hand-compiled city list and a per-document consistency pass. On DataNivra's internal synthetic benchmarks it still misses values it has no rule or context for, such as a bare first name in narrative text or a city in unusual phrasing, and those values stay in the masked text. The residual re-scan uses the same detectors, so it cannot see what they miss and cannot prove absence. Expect lower recall on real text (other languages, formats or OCR noise) than on synthetic corpora, and review samples of masked output in your own environment before relying on it.
  • Over-masking. Capitalised phrases after name cues, product names, order numbers, other dates on a date-of-birth line and some address-like phrases can be replaced even when they are not personal data.
  • Formats. PDF and legacy Office formats are not parsed. A PDF extractor is an optional extra awaiting owner approval.
  • Output shape. Masked documents are delivered as the documents table (one row per part), not re-rendered as the original files.
  • Dates of birth in a non-ISO format become a [DATE_OF_BIRTH] placeholder instead of a shifted date.
  • Cost. Entity counts on the job path come from one extra detection pass over the set, inside your network.
  • Models. Detector adapters and model clients are interfaces only; none is enabled or shipped.
  • Benchmarks. DataNivra publishes no accuracy figures for this capability. The internal benchmark methodology is documented for engineering use only.

← All documentation