Learning center

Narrated walkthroughs

Short guided tours you can listen to or read. Nothing plays until you press Play, and every walkthrough has captions and a full transcript.

Narration is produced by your browser’s built-in speech voice on your own device; no audio is downloaded and no voice is cloned. Your speed, mute and caption choices are remembered in this browser only.

What is TDM?

A two-minute introduction to test data management and why it matters. 2:00

Keyboard: Space or K play/pause, ← → seek 5 s, M mute, C captions.

Transcript
  1. Welcome. In the next two minutes, we will answer a simple question. What is test data management, and why should your team care about it?
  2. Every application is tested against data. When that data is thin or unrealistic, tests pass even though real defects are hiding. A claim that is denied and then resubmitted, or an account with a reversed payment, never shows up. So the bug ships.
  3. The quick fix is to copy production into a test environment. That gives realistic data, but it also spreads personal, health and financial information into places with weaker controls. Every copy is one more thing to protect.
  4. Test data management, or TDM, resolves that tension. It is the discipline of delivering the right data, in the right shape and the right privacy state, to every non-production environment, exactly when a team needs it. And removing it when they are done.
  5. A typical lifecycle has a handful of steps. A team requests a dataset. Sensitive fields are discovered and classified. A smaller subset is selected around a business entity, such as a member or a customer. Sensitive values are masked, or replaced with synthetic ones. The result is checked and certified, then provisioned into DEV, QA, SIT, UAT or a performance environment. Later it is refreshed or retired.
  6. Done well, TDM gives testers realistic, consistent data quickly. It gives security teams an audit trail they can follow. And it keeps production data where it belongs.
  7. Who needs it? Developers who want fast, small datasets. Testers who need realistic scenarios. And security teams who want to know where sensitive data lives.
  8. DataNivra is built around that idea. Realistic test data. Production data stays home. To go deeper, open the tutorial that accompanies this walkthrough, or try the synthetic interactive demo.

Voice: Your browser's default speech voice (rendered on your device). Original script © DataNivra. No recorded or third-party audio; speech is synthesized locally by the visitor’s browser.

Read the related tutorial: What is Test Data Management?

How DataNivra keeps production data in your environment

A three-minute tour of the boundary between the DataNivra control plane and your environment. 3:00

Keyboard: Space or K play/pause, ← → seek 5 s, M mute, C captions.

Transcript
  1. In this walkthrough, we will look at one promise and how it is kept. Production data stays in your environment. Let us see what that means in practice.
  2. DataNivra is split into two parts. The control plane is the service we host. It holds your organisation, users and roles, policies and approvals, dataset requests, job status, schedules and the audit trail. The data plane runs inside your own network. It is made up of the DataNivra agent, the test data engine and the industry packs.
  3. Every operation that touches a row happens in the data plane. Connecting to sources, profiling, classification, subsetting, masking, synthesis, validation, certification and provisioning all run next to your data. The hosted service does not even contain the code that processes rows.
  4. The agent connects outward to us, over HTTPS. We never connect inward to you. There are no inbound ports to open and no database to expose. The agent asks for work, receives a time-limited lease for one command, checks it, and runs it locally.
  5. Credentials never travel. Your source and target passwords, and your masking keys, stay in your own secret store. DataNivra configuration only holds a secret reference, such as the name of a vault entry. The agent resolves the value on its own side.
  6. Work is also bounded in time. A lease expires if the agent stops renewing it, and the job returns to the queue. Each command carries its own identifier, so it runs at most once. If the control plane cannot be reached, the agent finishes or stops in-flight work safely, never publishes uncertified output, and waits for a valid lease before starting anything new.
  7. So what does cross the boundary? Only four kinds of information. Control metadata, like job identifiers, states and column names. Aggregate metrics, like row counts and durations. Evidence references, like a checksum and the location of a report that stays with you. And secret references. Raw rows, samples, query results and credentials are prohibited.
  8. That rule is enforced in several layers. Message contracts refuse unclassified or prohibited fields. The agent has an egress guard that inspects every outbound message and blocks anything that looks like raw data. And the control plane validates again on arrival, without echoing what it rejects.
  9. We call this zero raw-production-data egress, and customer-resident processing. Notice that it is not zero copy. The agent does create subsets and masked datasets, but only inside your boundary.
  10. To see the boundary in action, try the synthetic interactive demo, where one panel shows everything the cloud receives.

Voice: Your browser's default speech voice (rendered on your device). Original script © DataNivra. No recorded or third-party audio; speech is synthesized locally by the visitor’s browser.

Read the related tutorial: Customer-resident data processing

From discovery to certified QA dataset

A four-minute walk through one job, from discovering sensitive fields to a certified dataset in QA. 4:00

Keyboard: Space or K play/pause, ← → seek 5 s, M mute, C captions.

Transcript
  1. In this walkthrough, we will follow a single job from start to finish. We begin with an invented health plan database, and we end with a certified dataset in a QA environment. Everything in the example is synthetic.
  2. Step one is connecting a source. An administrator registers the database by secret reference. The agent, running inside the customer network, resolves the credential locally and connects read-only. The control plane only learns the source identifier and the reference name.
  3. Step two is discovery. The agent reads schema metadata and profiles columns locally. It then classifies each column. First and last names and email addresses are direct identifiers. Date of birth is a quasi-identifier. Diagnosis codes and service dates are health information. Billed amounts are financial. What goes up to the console is a list of column names, classes and counts. No sample values leave.
  4. Step three is review. A data owner looks at the findings in the console and approves a classification policy and a masking policy. Each policy is versioned, and the approval is recorded in the audit trail. Separation of duties means the person who writes a policy is not the only one who approves it.
  5. Step four is the request. A tester asks for a QA dataset built around a business entity, say members, starting from a modest number of records. The request names the masking policy version, the target environment, the refresh schedule and the retention period.
  6. Behind the scenes, the control plane turns that request into a declarative command. It names the job, the subset definition and the policy versions to use, by reference. It does not contain any data. The agent picks the command up the next time it asks for work, and accepts a lease for it. From here on, all of the heavy lifting happens inside the customer network, and the console only sees progress events and counts.
  7. Step five is subsetting. The agent leases the job and first re-verifies the policy checksum. Then it selects the starting members and follows the relationships. It brings their coverage and claims, and every provider those claims refer to. This is referential closure. When it is done, no key in the subset points at a missing record.
  8. Step six is masking. Names and emails are replaced with realistic substitutes. Member identifiers become keyed pseudonyms. Because masking is deterministic, the same member identifier is replaced the same way in every table, so joins keep working. Dates are shifted consistently, and amounts are slightly perturbed.
  9. Step seven is validation and certification. The agent runs a set of gates. Coverage checks that the dataset has what was requested. Masking checks that no sensitive column survived unchanged. Referential integrity checks for dangling references. Schema checks the structure against what was approved. Quality rules and provenance checks finish the job, and a checksum is recorded.
  10. Certification fails closed. If any gate fails, the dataset is blocked, and a failed dataset is never provisioned. You fix the policy and run again. Masked is not the same as certified.
  11. Step eight is provisioning. The certified dataset is loaded into QA. The certification report stays in the customer environment. The console records the status, the aggregate counts and an evidence reference made of a location and a checksum.
  12. Finally, a refresh schedule rebuilds and re-certifies the dataset on a regular cadence, and retention rules retire it when it is no longer needed. You can try these same steps yourself in the synthetic interactive demo.

Voice: Your browser's default speech voice (rendered on your device). Original script © DataNivra. No recorded or third-party audio; speech is synthesized locally by the visitor’s browser.

Read the related tutorial: Dataset certification

Healthcare pack tour

A three-minute tour of the healthcare and life sciences industry pack. 3:00

Keyboard: Space or K play/pause, ← → seek 5 s, M mute, C captions.

Transcript
  1. Welcome to a short tour of the healthcare and life sciences pack. Industry packs are versioned plugins. They run inside your environment, alongside the agent and the engine, and they add domain knowledge that a generic tool does not have.
  2. The healthcare pack starts with an entity model. A member, or patient, has coverage under a plan. The member has encounters with providers. Claims refer back to encounters, providers and coverage. Knowing these relationships lets the engine build subsets where every claim still has its member, its provider and its coverage.
  3. Next come detection rules. The pack helps classify the columns you would expect. Names, addresses, birth dates and contact details. Member numbers, medical record numbers and other health plan identifiers. Diagnosis, procedure and medication codes. And free-text notes, which often contain the most sensitive detail of all.
  4. Then there are masking presets. You review them and approve them as your own policy versions. Dates are shifted, but all of one member’s dates move by the same amount, so the time between an admission and a discharge, or between two claims, is preserved. Member identifiers become consistent pseudonyms, so the same member lines up across eligibility, claims and encounter systems. Free-text notes are suppressed or replaced with synthetic text.
  5. The pack also includes synthetic scenarios, for cases that production rarely contains. For example, an invented member with several coverage periods and a mid-year plan change. Or a claim that is denied, resubmitted and then adjusted. These let testers exercise difficult paths without waiting for them to appear in real data.
  6. All of this runs inside your environment. The pack itself contains rules, templates and scenarios, never patient records. When the agent reports progress to the control plane, it sends the names of the columns it classified, counts of how many rules were applied, and gate outcomes. It never sends a diagnosis, a name or a member number, masked or not.
  7. Certification applies domain checks too. After masking, the gates confirm that identifiers were transformed, that relationships are intact, and that no protected health information remains in fields the policy says must be masked. If a check fails, the dataset is not provisioned.
  8. A word on regulation. This pack supports organisations running HIPAA-related privacy programmes, by keeping protected health information inside their environment and producing certification evidence. It does not, by itself, make any system meet a regulation. That remains a decision for your own privacy and compliance teams.
  9. To learn more, read the healthcare test data tutorial, or visit the healthcare pack page.

Voice: Your browser's default speech voice (rendered on your device). Original script © DataNivra. No recorded or third-party audio; speech is synthesized locally by the visitor’s browser.

Read the related tutorial: Healthcare test data

Financial-services pack tour

A three-minute tour of the financial services industry pack. 3:00

Keyboard: Space or K play/pause, ← → seek 5 s, M mute, C captions.

Transcript
  1. Welcome to a short tour of the financial services pack. Like every industry pack, it is a versioned plugin that runs in your environment, next to the agent and the engine. It adds banking and payments knowledge on top of the core platform.
  2. The pack begins with an entity model. A customer, or party, owns accounts. Accounts have cards, transactions and statements. Some accounts are joint, so one account can belong to more than one customer. The engine uses these relationships to build subsets where every transaction still has its account, and every account still has its owners.
  3. Detection rules help classify the sensitive columns. Names, addresses, national identifiers and contact details. Account numbers and international bank account style identifiers. Payment card numbers, which are recognised using checksum validation rather than just their length. And balances and transaction descriptions, which can reveal a great deal about a person.
  4. Masking presets are designed so that applications keep working. Card and account numbers are replaced with format-preserving substitutes that still pass checksum validation, so the application under test accepts them. Masked keys stay consistent across the core banking, card and statement systems. And balance-preserving transformations keep statements reconciling, so opening balance, transactions and closing balance still add up.
  5. Synthetic scenarios cover cases that are hard to find on demand. An invented account with an overdraft, a reversal and a chargeback. A customer who holds both joint and sole accounts. A month-end statement run where every balance must reconcile. Testers can generate these whenever they need them.
  6. As with every pack, the processing stays in your environment. The pack ships rules, templates and scenarios, not customer records. When the agent reports to the control plane, it sends column names, classes, rule counts and gate outcomes. Card numbers, balances and transaction descriptions never cross the boundary, whether masked or not. And the masking keys stay in your own secret store, referenced only by name.
  7. Certification then checks the result. The gates confirm that no card or account number survived unchanged, that relationships are intact, and that reconciliation rules still hold. If any gate fails, the dataset is blocked and never provisioned.
  8. About regulation. This pack supports PCI DSS oriented and privacy programmes, by keeping cardholder and customer data inside your environment and producing evidence. It does not certify any system against any standard. Assessments remain the job of your own teams and assessors.
  9. To learn more, read the financial services test data tutorial, or visit the financial services pack page.

Voice: Your browser's default speech voice (rendered on your device). Original script © DataNivra. No recorded or third-party audio; speech is synthesized locally by the visitor’s browser.

Read the related tutorial: Financial-services test data