Guide
Banks test some of the most interconnected systems in any industry: a core banking platform, payment hubs, card processing, online channels, fraud detection and KYC/AML screening, often across mainframe and cloud. Each system holds personal and financial data, and each test cycle needs data that behaves like production without being production. This guide covers what banking test data needs and how DataNivra approaches it. For the broader financial-services view, see financial-services test data.
Why banking test data is hard
- Every identifier has a format and often a check digit. An IBAN carries two mod-97 check digits, a card number (PAN) has a Luhn check digit and an issuer prefix, a routing or sort code identifies a real institution. Masked values that break these rules are rejected by validation logic before the test even reaches the code under test.
- Identities cross many systems. The same customer appears in the core ledger, the card platform, the CRM, the payment hub and the data warehouse, sometimes under different identifiers. A test of a payment flow needs every system to agree on who the customer is.
- Balances must reconcile. Transactions, holds and ledger balances are related by arithmetic, not only by keys. A subset that keeps accounts but drops some of their postings produces balances that never reconcile.
- Screening needs names, not only numbers. Sanctions and PEP screening, name matching and adverse-media checks are tested with names. Masked names that are random strings do not exercise fuzzy matching; realistic synthetic names do.
- The interesting cases are rare. Structuring patterns, chargebacks, dormant accounts reactivated, cross-border payments with unusual currencies, returned direct debits — these are exactly the cases a random sample misses.
What good banking test data needs
- Format-preserving masking with valid check digits. Masked IBANs keep the country code and length and carry recomputed check digits; masked card numbers stay in reserved test ranges with a valid Luhn digit, so test cards can never be real cards.
- Deterministic identities across systems. The same customer, account and card receive the same pseudonyms in the core system, the card platform and the warehouse. See cross-system deterministic masking.
- Entity-aware subsets that keep ledgers whole. Selecting a customer brings their accounts, every posting that makes up the balances, their cards and their payments — across systems — as described in cross-system test-data subsetting.
- Synthetic scenarios for the rare and the negative. Clearly labelled synthetic records for fraud-like patterns, screening hits, boundary amounts and invalid inputs, so negative tests do not depend on finding the right production case.
- Evidence that a dataset was certified. Masking coverage, integrity and the policy version recorded for each dataset, so a test environment's data can be explained to risk and audit teams later.
Core banking and ledgers
Core systems are frequently mainframe-based, with account and customer master files exported nightly to distributed systems. Keep the identifiers consistent between the mainframe extract and the modern databases it feeds, and include complete posting histories for selected accounts. The mainframe guide explains EBCDIC records, copybooks and packed-decimal amounts, and the current status of mainframe support, which is not shipped yet.
Payments and cards
Payment messages reference debtor and creditor accounts, agents and remittance information; card transactions reference cards, merchants and authorisations. Mask the account and card identifiers deterministically so a payment's debtor account still matches an account in the ledger, and keep merchant and currency reference data intact. Remittance and free-text fields can contain names and account numbers typed by people, so classify them and replace them rather than passing them through.
KYC, AML and screening
Screening and monitoring rules are best tested with synthetic customers designed to hit them: a name close to a watch-list entry, a customer whose transactions sit just under a reporting threshold, a sudden change in transaction geography. Label these records as synthetic so they are never mistaken for real alerts or real distributions.
How DataNivra approaches it
The financial-services industry pack provides classification rules, masking strategies (including card numbers generated in reserved test ranges with valid Luhn digits), relationship templates and synthetic scenarios. Masking runs in the agent inside your network with keys held in your secret store; deterministic, keyed strategies keep identities consistent across systems. Certification refuses a dataset with incomplete masking or broken relationships, and nothing is provisioned until it passes. DataNivra makes no compliance claim on your behalf: the pack helps you produce safer test data, it does not by itself make any system compliant.
Check which of your systems can be read today on the integrations page, estimate dataset sizes with the free tools, read the documentation, or start free with a synthetic banking estate.