Guide
Core banking, insurance policy administration, payroll and government benefit systems still run on mainframes. Their data rarely stays there: nightly batch jobs export records that feed distributed databases, warehouses and partner files. Testing an end-to-end flow therefore needs mainframe records and modern-system records that describe the same people and accounts — masked consistently on both sides. This guide explains what makes mainframe test data different and what DataNivra offers today.
Current status: mainframe files are in preview
DataNivra ships a preview mainframe file connector (MAINFRAME_FILE). It reads unloaded extract files — the sequential files produced by batch jobs or IDCAMS REPRO — from a read-only landing directory inside your network, decodes them with your COBOL copybooks and feeds the decoded records into the same classification, masking, subsetting and certification workflow as every other source. The mainframe files integration page always shows the live status, the tested formats and every known limitation.
Preview limitations you should know before planning:
- It was verified with synthetic EBCDIC datasets written by DataNivra's own record encoder; it has not yet been verified against unload files produced on a z/OS system.
- It reads extract files only. It does not connect to z/OS, VSAM or Db2 for z/OS directly; moving the unloads to the landing directory is part of your batch process.
REDEFINES,OCCURS DEPENDING ONand floating-point (COMP-1,COMP-2) items are refused with a reason code instead of being guessed.- Variable-length files are read in one pass up to a configurable size limit.
- Masked output is written as ordinary tables (for example Parquet in a directory target); writing it back in the original record layout is not implemented yet.
What makes mainframe records different
- EBCDIC encoding. Text fields are encoded in an EBCDIC code page such as IBM-037 (US) or IBM-1047 (Open Systems), not ASCII or UTF-8. Reading a file with the wrong code page silently corrupts names and addresses, and national code pages differ in where they put characters such as
[,]and currency symbols. - Copybooks define the layout. A COBOL copybook describes each field's position and type:
PIC X(30)for a 30-character text field,PIC 9(7)for zoned decimal digits,PIC S9(7)V99 COMP-3for a signed packed-decimal amount with two implied decimal places. Without the copybook, a record is just bytes. - Packed decimal (COMP-3). Two digits are packed into each byte and the last half-byte holds the sign (
Cpositive,Dnegative,Funsigned). A masking step that treats these bytes as text destroys the amount. - Record formats. Fixed-block files have records of one length; variable-block files prefix each record with a four-byte record descriptor word.
REDEFINESlets one area mean different things depending on a record-type field, andOCCURS DEPENDING ONmakes the length of a repeating group depend on a count stored earlier in the record. - VSAM and sequential exports. Keyed VSAM data sets are typically exported to sequential (QSAM) files with utilities such as IDCAMS
REPRObefore anyone off the mainframe can read them.
What good mainframe test data needs
- Decode with the right copybook and code page, inside your network, into typed columns.
- Classify field by field. Customer names, national identifiers, account numbers and dates of birth sit next to amounts and flags in the same record.
- Mask without changing the layout. A masked
PIC X(10)is still exactly ten characters; a maskedS9(7)V99 COMP-3is still a valid packed decimal with the same scale and sign convention. Otherwise the test region's batch jobs reject the file. - Keep identities consistent with modern systems. The customer number in the mainframe extract is usually the same identifier as
customer_idin a distributed database. Both must receive the same pseudonym, or the end-to-end test breaks at the first join. See cross-system deterministic masking. - Subset by entity across the estate. Selecting a customer should bring their mainframe account records and their records in modern systems, as described in cross-system test-data subsetting.
- Re-encode for the target. If the masked data goes back to a mainframe test region, it has to be written in the original layout and code page.
What works today
- Unload the data sets you need to sequential files and deliver them, with their copybooks, to a landing directory that is mounted read-only into the agent. The agent proves from filesystem permissions that it cannot write there before it reads anything.
- Register a
MAINFRAME_FILEsource that declares each data set: a table name you choose, the record file, the copybook, the record format (F,FB,VorVB), the code page (cp037by default;cp500,cp1140,cp273,cp1026andcp875are supported) and which numeric fields hold dates (YYYYMMDD,YYYY-MM-DDorYYMMDD). Table names are never taken from data set names, which often embed data. - The connector decodes text, zoned decimal with implied decimal places, packed decimal (
COMP-3) and binary (COMP,COMP-4,COMP-5) fields, flattens group items and fixedOCCURSgroups into columns, and refuses any record whose length does not match the copybook. Invalid packed or zoned digits fail withMAINFRAME_INVALID_NUMERICrather than producing a half-decoded value. - Files carry no keys, so declare the relationships between mainframe tables and your other sources through an industry pack or explicit mappings. Use the same masking keys and rules for identifiers that also appear in distributed systems so pseudonyms line up across the estate.
- Discover, classify, subset, mask and certify as for any other source, then provision the certified dataset to a directory target.
If your extracts use features the preview refuses, the earlier path still works: convert them to Parquet or CSV with your own copybook-aware tooling inside your network and read them with the Local files connector. Treat any converted extract as production data and delete it once the certified dataset exists.
How DataNivra approaches it
Decoding happens inside the agent, in your network: records are turned into typed Arrow columns and never leave the host. Only metadata (table and field names from your copybooks, declared types), aggregate profiles and evidence cross the boundary. Mainframe fields join the same entity graph, masking keys and certification gates as every modern source, so a customer number masked in the mainframe extract and in a distributed database receives the same pseudonym.
Next steps
- Read the mainframe category overview and the banking test data guide.
- Plan the scope of a first dataset with the free tools.
- Evaluate the workflow on synthetic data: start free or open the documentation.