Learning center

Data discovery and classification

How to find sensitive fields across databases and files, classify them, and turn findings into approved policy without exposing values.

Tutorial

You cannot protect what you have not found. Sensitive-data discovery is the work of locating personal, health, financial and confidential information across tables, columns and files. Classification labels what was found so that policy can decide how to treat it.

In real estates the obvious columns — ssn, email — are the easy part. The hard part is notes, ref_2 and the column someone repurposed years ago.

Discovery and classificationInside your environment the agent connects read-only, reads schema metadata, profiles columns locally and classifies sensitive fields. Only column names, classes and counts are sent up. In the cloud, reviewers approve findings and publish a classification policy.Your environmentConnect read-onlyRead schema metadataProfile columnslocallyClassify sensitivefieldsDataNivra CloudReview findingsApprove classificationpolicyVersion & auditFindings: column names, classes, countsNo sample values leave your environment
Discovery and classification. Inside your environment the agent connects read-only, reads schema metadata, profiles columns locally and classifies sensitive fields. Only column names, classes and counts are sent up. In the cloud, reviewers approve findings and publish a classification policy.
Text description
  1. Your environment: Connect read-only → Read schema metadata → Profile columns locally → Classify sensitive fields

    Connection: Findings: column names, classes, counts — No sample values leave your environment

  2. DataNivra Cloud: Review findings → Approve classification policy → Version & audit

Why it matters

Masking and subsetting are only as good as the inventory beneath them. A missed column means real values flow into test environments untouched, no matter how strong the masking policy is for the columns you did find.

Classification also has to capture risk that is not obvious from a single field. A Quasi-identifier such as postcode or birth date is harmless alone but identifying in combination. And discovery must be repeatable: schemas change, so a one-off spreadsheet inventory goes stale within a release or two.

Example

A synthetic members table profiled locally might produce findings like these:

ColumnSignalsProposed class
full_namename pattern, high distinct ratioDirect identifier
dobdate type, plausible birth rangeQuasi-identifier
postcodepostal formatQuasi-identifier
plan_codelow cardinality, short codesNot sensitive
noteslong free text, name-like tokensFree text: suppress or synthesize

The column names and signals are invented. Note what the findings report contains: column names, classes and counts — never the sample values that led to the conclusion.

How DataNivra approaches it

The agent connects to sources read-only and does the profiling inside your data plane. Detection combines schema metadata, data-type checks, pattern and checksum detectors, and rules from industry packs. Any local samples used for profiling are treated as prohibited raw data: they stay in the job sandbox and are cleaned up.

What reaches the control plane is a findings report of names, proposed classes and aggregate counts. Data owners review the findings in the console, adjust them, and approve a versioned classification policy with separation of duties. Rediscovery on later runs highlights new or changed columns so that drift is caught before data is published. See Customer-resident architecture for the boundary details.

Key takeaways

  • Discovery is the foundation for every other TDM control.
  • Look beyond obvious columns: free text and quasi-identifiers matter.
  • Findings can be shared as metadata; the values behind them should not be.