Tutorial
You cannot protect what you have not found. Sensitive-data discovery is the work of locating personal, health, financial and confidential information across tables, columns and files. Classification labels what was found so that policy can decide how to treat it.
In real estates the obvious columns — ssn, email — are the easy part. The hard part is notes, ref_2 and the column someone repurposed years ago.
Text description
- Your environment: Connect read-only → Read schema metadata → Profile columns locally → Classify sensitive fields
Connection: Findings: column names, classes, counts — No sample values leave your environment
- DataNivra Cloud: Review findings → Approve classification policy → Version & audit
Why it matters
Masking and subsetting are only as good as the inventory beneath them. A missed column means real values flow into test environments untouched, no matter how strong the masking policy is for the columns you did find.
Classification also has to capture risk that is not obvious from a single field. A Quasi-identifier such as postcode or birth date is harmless alone but identifying in combination. And discovery must be repeatable: schemas change, so a one-off spreadsheet inventory goes stale within a release or two.
Example
A synthetic members table profiled locally might produce findings like these:
| Column | Signals | Proposed class |
|---|---|---|
full_name | name pattern, high distinct ratio | Direct identifier |
dob | date type, plausible birth range | Quasi-identifier |
postcode | postal format | Quasi-identifier |
plan_code | low cardinality, short codes | Not sensitive |
notes | long free text, name-like tokens | Free text: suppress or synthesize |
The column names and signals are invented. Note what the findings report contains: column names, classes and counts — never the sample values that led to the conclusion.
How DataNivra approaches it
The agent connects to sources read-only and does the profiling inside your data plane. Detection combines schema metadata, data-type checks, pattern and checksum detectors, and rules from industry packs. Any local samples used for profiling are treated as prohibited raw data: they stay in the job sandbox and are cleaned up.
What reaches the control plane is a findings report of names, proposed classes and aggregate counts. Data owners review the findings in the console, adjust them, and approve a versioned classification policy with separation of duties. Rediscovery on later runs highlights new or changed columns so that drift is caught before data is published. See Customer-resident architecture for the boundary details.
Key takeaways
- Discovery is the foundation for every other TDM control.
- Look beyond obvious columns: free text and quasi-identifiers matter.
- Findings can be shared as metadata; the values behind them should not be.