Skip to content

White papers

Version 1.8 · Public 3 September 2026

The Oracle You Already Own

Legacy modernization as an evidence problem

Pierre Boutquin

For CIOs, CTOs, heads of engineering, and audit stakeholders responsible for business-critical systems that cannot be allowed to fail.

Short form: Executive brief

First published
25 August 2026
Readability
Flesch–Kincaid Grade 11–14

Five-minute reading path

Read this table only. The links are optional evidence dives.

Minute Executive read Optional evidence dive
0–1 Name the programme. A conversion preserves behaviour on a new platform. A re-implementation changes behaviour deliberately. A replacement changes the capability. Equivalence is the objective only for the first. A sponsor who cannot name three post-cut-over changes, owners, and dates has not yet shown why a like-for-like conversion creates value. That screen does not replace a quantified investment case. §1.1–1.4
1–2 Approve a bounded evidence pilot, not conversion. Its charter fixes scope, duration, spend, accountable executive, risk appetite, data authority, regulatory applicability, and excluded live effects. The pilot captures a versioned legacy oracle, calibrates comprehension blind, declares budgets before results, and performs non-mutating replay. Contract the right to reshape, pivot, or stop without weakening a failed gate. §5 and §12
2–3 Ask three questions in order, backed by six measurements. First: Is it complete? Oracle coverage and comprehension calibration account for the declared evidence population and expose every gap. If required evidence is missing, the answer is no. Second: Is it consistent? Behaviour-and-service divergence plus data-and-security posture expose functional, operational, data, and control differences without averaging them together. Third: Where does it say that? Gate integrity and drift/freshness make the accepted result traceable and current. Every value carries a denominator. Unknown evidence is UNEVIDENCED, never a favourable zero. For regulated entities, §2 adds an applicability gate before this sequence and an obligation-by-obligation compliance determination after it. §2 and §7
3–4 Make differences decidable. The legacy system witnesses what production did. It does not decide what the business should do. Every deliberate departure belongs in an intent register with a person, rationale, evidence, date, and change-budget consumption. Every non-idempotent or external effect belongs in a one-way-door register. Reversible claims require tested compensation. The first irreversible external write is a release-specific human gate. §9 and §11
4–5 Know the claim boundary. No evidence method turns an unknown population into a favourable zero. The public examples show how to name a coverage gap and assign responsibility, not client outcomes or implementation schemas. The method reports measured divergence, visible gaps, and accountable decisions, not universal identity or zero risk. §13 and selected public examples

Contents

  1. What are you actually buying?
  2. Regulatory scope and independent review
  3. The claim nobody should make
  4. What actually breaks
  5. The oracle you already own
  6. The moving target
  7. Three questions, six measurements
  8. Composition risk
  9. Divergence triage and intent
  10. Legacy experts as a control
  11. The one-way door
  12. Pilot outcomes and off-ramps
  13. Claim boundaries
  14. The cost of evidence
  15. Decision synthesis
  16. Appendix A—numbered evidence register
  17. Appendix B—selected sample evidence artifacts
  18. Appendix C—glossary
  19. Appendix D—evidence gaps and excluded claims

Before anything else

A modernization pipeline can look safe on a methodology diagram. The boxes and arrows omit denominators, missed behaviours, stale evidence, and the person who accepted a difference. Those omissions determine whether the general ledger still balances in March.

Evaluation starts with the investment case. What does a like-for-like conversion make possible, and who will use that option? The next question is operational: how will the programme detect a different answer from the new system? Once a difference appears, someone with business authority must decide which answer should govern future behaviour.

Detection receives most of the industry's attention because a methodology diagram appears to answer it. The investment question comes earlier. A like-for-like conversion creates no new business capability by construction, so its return depends on changes made afterward. The business decision comes later, when the programme must choose between current intent and rules written by people who left in 2004.

Each answer needs an operating mechanism and an accountable person. The investment case and evidence gates also need numbers. Without those elements, work may be underway, but it is not governed.

The evidence work has an order. Declare and account for the full in-scope population, then close its material gaps before making a consistency claim. Resolve every inconsistency before publishing a canonical result. Then make the result traceable to the evidence and authority that produced it. A team can build the controls in parallel, but it cannot pass a later question while an earlier one remains open.

Regulated work adds a gate on each side. Applicability determines which obligations shape the declared scope before completeness begins. A qualified human decision-maker assesses compliance only after the inner evidence sequence closes. That assessment cannot repair an incomplete population, an unresolved inconsistency, or a missing deed.

1. What are you actually buying?

A like-for-like conversion produces no new business capability. That is not a criticism. It is the definition: same behaviour, different platform. If it produced new capability, it would not be like-for-like. The paper's equivalence testing would then measure the wrong thing.

Before proceeding, ask what the organisation buys if behaviour remains identical.

1.1 Three different programmes wear the same word

"Modernization" names three programmes with different objectives, risk profiles, and uses for the legacy system.

What changes What the legacy system is Equivalence testing is
Conversion (like-for-like) Platform only The target: match it The whole job
Re-implementation Behaviour, deliberately A reference: identify what changes A comparison tool, not a gate
Replacement (package, rebuild) The capability itself An inventory: identify what disappears Mostly irrelevant

This paper addresses conversion. If the actual intent is re-implementation, proving equivalence would establish fidelity to behaviour that the programme plans to change. That is an expensive guarantee with little value. For replacement, the legacy system supplies a capability inventory rather than a target specification.

Programmes get into trouble by starting as one and becoming another without a decision. §9 examines how that happens one divergence at a time.

1.2 What a conversion is worth

If behaviour is unchanged, the value is not in the system. It is in what becomes possible:

  • Change velocity: the ability to make changes that are currently too slow, too risky, or refused outright
  • Hiring: access to a labour pool that is not retiring
  • Integration: an interface surface that modern systems can reach
  • Vendor and hardware exit: removal of a dependency you cannot negotiate against
  • Reporting and audit access: access to data that is currently trapped

Each item is an option rather than a realised benefit. A conversion that stops at cut-over has created the option but has not exercised it. The return appears in later changes, which belong in the business case before work starts.

Ask the sponsor to name the changes planned for the first twelve months after cut-over that cannot be made today. Each change needs an owner and a date. If no such list exists, the programme has not shown how the conversion will create value. Stopping at that point is cheaper than reaching the same finding at a steering committee eighteen months later.

1.3 The obstacles that are not about code

The feasibility constraints arrive before engineering gets a vote.

Running both systems costs more than running one. Mainframe software charges may not fall in proportion to the workload removed. Under IBM's traditional sub-capacity pricing, charges track the highest observed four-hour rolling average (R4HA) for the relevant logical partitions (LPARs). They do not track total monthly consumption. [1] Therefore, migrating a workload outside the monthly peak window produces no reduction in that charge.

Model the business case to the last workload in the peak window. Use the monthly peak R4HA, not an average measured in millions of instructions per second (MIPS).

A strangler programme that delivers new functionality as it goes can stop partway and still pay. Fowler's original example did exactly that. [2] A pure like-for-like conversion does not. An incremental programme may bank value early. A stalled conversion has bought a second platform and retired none of the first. The transitional architecture that allows coexistence is real engineering and belongs in the estimate. That estimating rule is this paper's inference, not Fowler's.

Some platform behaviour cannot be reproduced exactly. Record that result as a finding and make the relevant semantics explicit:

Legacy semantic Modernization risk Primary reference
COBOL decimal arithmetic and ROUNDED; COMP-3 packed-decimal storage; the ARITH(EXTEND) compiler option for extended precision; Extended Binary Coded Decimal Interchange Code (EBCDIC) collation; REDEFINES storage overlays Changed rounding, overflow, ordering, or record interpretation IBM COBOL 6.4 [3]
Customer Information Control System (CICS) pseudo-conversation and communication area (COMMAREA) state Different task, lock, restart, and session boundaries IBM CICS Redpaper [4]
Virtual Storage Access Method (VSAM) key, relative-record, and relative-byte access Behaviour coupled to physical layout IBM VSAM Redbook [5]

These are not academic edge cases. A mixed alphanumeric key can sort differently after an EBCDIC-to-UTF-8 move. A monetary value exact in packed decimal can change under binary floating point. A CICS task can deliberately end while carrying state in a COMMAREA. Choosing a session framework cannot “preserve” that task. It replaces the interaction model. Code that addresses VSAM data by relative byte address makes physical layout observable behaviour.

The programme therefore needs a platform-semantics inventory before estimation. It records each legacy semantic, where it appears, whether the target reproduces or replaces it, affected evidence paths, and the business owner for the decision. Without it, “like-for-like” silently acquires implementation exceptions and becomes re-implementation by accumulation.

Surface these in the first weeks. Each is an intent-register decision, not a bug for an engineer to settle quietly. Also price the scarce overlap of legacy and modern skills. Price the downstream partners whose formats, timing, and change calendars are not yours.

1.4 When not to do this

Modernization is not always the appropriate choice. Four conditions should stop or redirect the work.

  • The business wants different behaviour. Then equivalence is the wrong objective. Do re-implementation deliberately, with the oracle as a reference rather than a target, and price it as what it is.
  • The capability is commodity. If the market sells the payroll, general-ledger, or CRM capability, matching thirty years of local idiosyncrasy may cost more than a package evaluation.
  • Nothing will change after cut-over. Per §1.2, the option is the product. No exercise plan, no return.
  • The system is genuinely stable and genuinely isolated. Some legacy systems are working, unchanging, and cheap. "It is old" is not a business case. Encapsulate it behind an interface and revisit in three years.

Industry taxonomies support this broader choice set. AWS's seven-strategy guidance treats retire, retain, rehost, relocate, repurchase, replatform, and refactor as distinct options, and is direct about the most invasive one:

"Refactor is not recommended for large migrations because it involves modernizing the application during the migration… For a large migration, refactor only when the other migration strategies are not an acceptable option." [6]

AWS has a commercial interest in migration volume, yet its guidance recommends against the highest-effort path for large estates. The practitioner study cited throughout this paper also records why organisations retain legacy systems: interviewees described them as business-critical, proven, reliable, and performant. [7]

2. If you are regulated, this is not only an engineering decision

For a bank or insurer, the applicable regime changes who signs and what the programme must record. It also determines whether independent review is mandatory. The programme must establish scope before it borrows a control from adjacent guidance.

The regulated wrapper

The three-question sequence is universal evidence discipline. Regulated entities add two outer gates rather than treating “compliance” as a synonym for traceability:

Order Question Required record Gate before proceeding
0 What applies? The programme's applicability register names the jurisdiction, instrument, legal force, entity, role, system and classification, obligation population, and effective date The programme classifies every candidate obligation as applicable or not applicable. A not-applicable disposition names its authority and rationale.
1–3 Complete → Consistent → Traceable The programme produces the six measurements, intent decisions, protected evidence, and current person-linked deed described in §7 The programme accounts for the declared evidence population, gives every material inconsistency an authorised disposition, and makes the result traceable.
4 Is it compliant? A qualified human decision-maker owns the obligation-by-obligation assessment The decision-maker classifies every applicable obligation as satisfied, not satisfied, or UNEVIDENCED. A positive conclusion is permitted only when none is not satisfied or UNEVIDENCED.

A traceable test result is not a compliance determination.

The compliant claim keeps its denominator beside it: compliant for the declared regulatory scope, as of the stated date. That scope names the entity, role, system, jurisdiction, instruments, obligation population, and effective date. A generated report is an input to qualified human judgment, not a standalone conformity assessment, legal opinion, or cut-over authorisation.

2.1 First decide whether the system is a model

OSFI E-23 is final but, as of this paper's status date, is not yet effective. It takes effect 1 May 2027. It will apply to applications that meet its three-part definition of a model. The definition requires:

  • data inputs
  • a processing component that uses theoretical, empirical, judgmental, statistical, or artificial intelligence and machine-learning (AI/ML) methods
  • useful results

It does not turn all software into a model. For an in-scope model, changes to algorithms, parameters, or supporting operational components trigger model review. That review must be independent of development. [8]

If the migrated system is in scope, “we intended to change nothing” is a testable claim, not an exemption. If it is not, the controls below may still be sound engineering without being an E-23 requirement.

2.2 Similar controls appear in different scopes

Jurisdiction Instrument Status and relevant scope
🇨🇦 Canada OSFI E-23, Principle 3.4 Final supervisory guideline for in-scope federally regulated financial institution (FRFI) models. It is effective 1 May 2027, not yet effective at the paper's status date. It requires independent model review. [8]
🇬🇧 UK PRA SS1/23, Principle 4.1(d)–(e) Supervisory statement for firms in scope. Validation is independent from development and model owners. [9]
🇨🇭 Switzerland FINMA 08/2024 §2.7 Supervisory communication reporting that separation was absent in some observed AI applications. Finding, not rule text. [10]
🇺🇸 United States SR 26-2 §§II–III, VI Interagency supervisory guidance, most relevant to banking organisations over US$30 billion and sometimes smaller banks with significant model risk. It covers traditional statistical/quantitative and non-generative, non-agentic AI models. Generative and agentic AI are excluded. It is non-binding guidance. Effective challenge needs sufficient independence and organisational influence. [11]
🇪🇺 Euro area ECB guide to internal models §§15, 19 Supervisory expectations for internal models. Validation staff are separate from development staff. [12]
🇪🇺 EU AI Act Arts 6, 17, 113, Annex IV §2(a) Primary legislation for providers of systems classified as high risk. Under the 2026 amendment, relevant Chapter III duties apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I product systems. [13]

These instruments do not create one universal modernization rule. Where independence binds, staffing it is a programme-design decision: a name added at final sign-off is not an independent review function.

2.3 For a covered high-risk system, development tooling belongs in technical documentation

For a provider of a covered high-risk AI system, Annex IV §2(a) requires technical documentation to describe “the methods and steps performed for the development of the AI system, including, where relevant, recourse to pre-trained systems or tools provided by third parties and how those were used, integrated or modified by the provider”. The named objects are pre-trained systems and third-party tools. The duty is qualified “where relevant”. It therefore reaches development tooling through the AI supply chain rather than as a general software-provenance rule. Under the amended transition dates, those Chapter III duties are not yet applicable at this paper's status date. [13] The provision supports a tool-and-provenance record when in scope. It does not expressly require an “AI-authored percentage.” Using an AI coding assistant does not by itself make the resulting software a high-risk AI system.

2.4 The independence requirement lands on the AI, and that part is an inference

The next control is an inference rather than a requirement from the instruments above.

If one agent writes both the migrated code and the tests that judge it, both artifacts descend from one interpretation of the legacy source. This paper treats that as a machine-separation failure.

This machine-separation control is a synthesis. The cited instruments address people, functions, covered systems, and organisational units. They do not address AI authorship of code and tests. The control should be presented for challenge rather than attributed to a regulator.

The mechanism behind the control predates language models. In the classic multi-version programming experiment, two universities independently wrote twenty-seven versions of one programme from the same specification. Researchers ran them against one million test cases. The versions were individually reliable. Yet more than one version failed on substantially more tests than independent failure predicts. Researchers rejected the independence assumption at the 99% confidence level. [43] Separate development from a shared specification did not deliver independent failure in 1986. The evaluation research below measures the same property in current tooling.

Evaluation research supports the mechanism but supplies no migration-specific result. ICML 2025 work measures language-model similarity through overlapping mistakes. It reports two findings relevant to this control. A model acting as judge scores similar models more favourably. Mistakes also become more similar as capability increases. [41] A separate evaluation study finds that a judge panel reduces single-judge bias only when it draws on disjoint model families. [42] Together, these results make independence a property of provenance rather than instance. A second agent from the same family changes the actor without decorrelating errors. The model studies evaluate generated text, and the multi-version experiment predates language models. None measures legacy-code migration. Applying them here is this paper's inference, not a demonstrated result. The inference is conservative because it argues for stronger separation than either study establishes.

2.5 Independent review has separate gates

The distinction matters because they fail independently:

What must be separate Failure looks like
Human separation The person who built the migration from the person who reviews it One engineer authors and merges. Review is a rubber stamp or absent.
Machine separation The agent that writes the code from the artifact that judges it Tests generated alongside the code, from the same reading of the legacy source (§4)

A programme can satisfy human separation while failing machine separation. For example, an independent person may approve a test suite that the AI produced from the same interpretation used to write the code. The reverse can also occur: judging evidence may have independent provenance and write protection, but the migration author may remain the only human reviewer.

Machine separation has two gates. Provenance separation asks who created, selected, transformed, normalised, and set thresholds for judging evidence. Access protection asks who could change or skip it afterward. Evidence produced by the migration author does not become independent when locked. Independently sourced evidence is not protected if the author can later relax it.

When a model did any of that work, the provenance record names its model family and version. It does not name only the actor who ran it. On the evidence in §2.4, another agent from the migration's model family does not provide separate provenance. A programme that reports the agent as independent has recorded the wrong variable. Execution of the legacy system produces evidence that no model authored. This paper recommends that separation wherever the programme can capture the behaviour (§5).

Measurement 5 reports provenance separation, access protection, and human separation as distinct results. An assertion that "review was independent" has no recoverable meaning unless it identifies the measured gate. A fourth result, demonstrated sensitivity, asks whether the gate can return red on a case built to fail it. Protected evidence has little value if the scoring path cannot reject anything (§7).

3. The claim nobody should make

The standard promise in this market is behavioural equivalence: the modernized system will behave identically to the legacy system.

That promise should not be made without a defined population and a measurement that can disprove it.

In 2026, IBM Research and Columbia released ScarfBench, a set of 204 directed enterprise-Java cross-framework migration tasks graded by executable oracles. A result had to compile, deploy in the target framework's containerised runtime, and pass behavioural tests. Across five coding agents, the strongest aggregate test-pass rates were 15.3% on focused-layer migrations and 12.2% on whole applications. One task was fully behaviourally equivalent. [14]

That is one fully equivalent task in the tested benchmark, not a base rate for legacy estates.

The benchmark keeps the language and Java Virtual Machine (JVM) ecosystem. However, it changes frameworks and runtime models across Spring, Jakarta EE, and Quarkus. It is directly relevant to cross-framework migration. It is not a COBOL-to-Java proxy.

A second measurement bounds it from a different, narrower task. Amazon Science's MigrationBench evaluated Java 8 → Java 17. It reported 71.67% first-attempt success (pass@1) for minimal migrations and 53.33% for maximal migrations on a 300-repository subset. [15]

Even within the Java ecosystem, a quarter to a half of repositories failed those tasks.

Low single-attempt rates invite an obvious response: generate many candidates and keep the best. Inference-scaling research measures when that works. The share of problems solved by at least one sample grows log-linearly across four orders of magnitude. That coverage converts into delivered performance where answers can be automatically verified. Without an automatic verifier, majority voting and reward models plateau beyond several hundred samples. [44] Repeated generation therefore buys equivalence only after the programme owns an executable check. The verifier is the scarce asset. §5 concerns the one already owned.

Universal, all-input identity is not a testable delivery claim. The programme can test observational equivalence over a declared population of inputs, states, parameters, and boundary effects. It can bound coverage, measure divergence, and expose gaps. Each paper claim can fail visibly.

The programme can verify a narrower claim. It reports the measured divergence rate against a budget approved before testing. The test replays production traffic through both systems. This exposes the number, population, and budget to challenge instead of hiding them inside a promise of identity.

4. What actually breaks

The most dangerous translation failure is delayed discovery: a rule present in the old system never reaches the new one, and no generated test notices.

The risk predates generative AI. Bisbal et al. describe the cut-over and coordination risks of incremental migration. The CMU Software Engineering Institute warned that legacy systems encapsulate business expertise. It gave no guarantee that a replacement would be as robust or functional. [16] [17]

AgentModernize (Ahmed & Galib) measured a 2026 version of the same failure mode and describes it this way:

"Implicit rules, edge-case handling, and cross-module constraints that keep production systems running are lost, and nobody notices until something fails in production."

On eight synthetic scenarios, the paper reports a 23.0% mean behavioural equivalence rate for its GPT-4o-mini feedback configuration, not a 23% error rate. The result is evidence that the authors' pipeline still struggled to produce equivalent implementations. It is not a real-world estate-wide failure baseline. [18]

The failure is silent because of the structure of the workflow:

   Legacy system
        │
        ▼   [AI reads the code]        ← if a rule is missed HERE …
   Extracted "business rules"
        │
        ├──────────────► new code       ← … it is absent here
        └──────────────► test suite     ← … and absent here too
                                           so the test cannot fail

A rule the analysis never extracted cannot appear in a test written from that analysis. The test suite and the code share a parent. They inherit the same blind spots, they fail in the same direction, and the gate goes green.

The testing literature calls the missing artifact a pseudo-oracle: an alternative implementation produced independently. Independence is essential because two implementations provide useful evidence only when they can fail differently. [19]

Independent production provides a weaker guarantee than it appears to. In the multi-version experiment in §2.4, teams wrote versions independently from one specification. Those versions still failed together far more often than independence predicts. [43] Therefore, an oracle's value rests on measured rather than assumed independence. This is why §5 turns to a witness whose provenance no reading of the source shares.

4.1 Two oracles, two jobs

“Oracle” is used for two different controls in this paper. They must not be collapsed:

Control What supplies the expected result Where it is used What it can establish
Recorded legacy oracle (golden-master evidence) Versioned observations captured from the running legacy system M1 coverage and M3 divergence What the legacy system did for the captured inputs, states, parameters, order, and boundary conditions
Independent calibration oracle/key A hidden, independently prepared rule/behaviour key. When an alternative implementation supplies it, that implementation is a pseudo-oracle. M2 comprehension calibration What the extraction method recovered or missed on the scored calibration scope

The recorded legacy oracle is an empirical witness, not ground truth. It cannot prove that historic behaviour was intended, lawful, desirable, or complete. The calibration oracle is independent only if its authors and evidence are separate from the extractor and candidate. Tests from the same machine reading as the migrated code satisfy neither control. They restate the same opinion as an assertion.

The programme already owns one system-level witness before it creates an independent calibration key. The running legacy system does not descend from the migration's interpretation.

5. The oracle you already own

The legacy system can serve as a specification for what it demonstrably did under captured conditions. Documentation may be stale, and people tend to remember the parts that caused incidents. The running system has executed the actual rules against production data without inheriting the migration's interpretation.

That independence makes it a uniquely useful empirical witness. It does not make the behaviour correct or the capture complete. The programme must separately validate the instrumentation, decoding, state alignment, and normalisation that turn execution into recorded evidence.

Its evidentiary value ends when the system is retired, so capture must happen while production behaviour remains observable.

Capture therefore precedes translation analysis.

5.1 Capture and replay have different control problems

A useful corpus must preserve enough context to distinguish a migration defect from capture loss, concurrency noise, or an ordering mistake. The implementation is platform-specific. “Use change data capture (CDC)” is not an architecture. Five control requirements apply regardless of the capture product:

  1. Observe without silently changing production. Select an authorised, platform-native observation point. Examples include logs, journals, broker taps, mirrored traffic, and application events. Benchmark its latency, locking, storage, failure, and recovery effects under representative load. An independent reviewer validates capture completeness and decoding. “Asynchronous” is not a synonym for “non-mutating” or “zero-risk.”
  2. Bind events to a consistent logical state. Record correlation keys, transaction boundaries, ordering domains, commit/effective times, delivery/retry semantics, isolation/read versions, and code/parameter versions. Bind the corpus to a state snapshot or hash and capture-log offset. Legacy-secondary and candidate executions start from that same declared state. Otherwise, the comparison measures different worlds.
  3. Preserve causality, not cosmetic order. Replay causally independent partitions in parallel. Preserve order where it is part of the contract. Compare sets or invariants only where order is explicitly irrelevant. Never sort an observable output merely to make two systems agree.
  4. Control non-determinism explicitly. Virtualise clocks, random sources, generated identifiers, and external responses where doing so is safe and faithful. Independently review and version every normalisation or injection rule. Residual variation is retained for event-level attribution against the legacy self-noise baseline, never placed in an undocumented ignore list. The primary/secondary/candidate pattern provides a practical model. [28]
  5. Isolate every replay effect. Route both legacy-secondary and candidate writes to shadow state or stubs and compare the attempted effects. Neither replay may post, settle, file, notify, or consume a live sequence. Scientist's safety warning is direct: experiments that change data are unsafe without isolation. [29]

This design deliberately does not prescribe Debezium, a VSAM log reader, container-level clock mocking, or any other product-specific topology. Those may be valid in a particular estate, but they require a platform assessment and cannot be inferred from “mainframe,” “CDC,” or “shadow run.”

5.2 Harvest before you translate

Instrument the legacy system before the team reads code for translation. Record its observable behaviour as a replayable corpus. The corpus contains inputs, outputs, state transitions, timing, and calls across system boundaries.

The corpus supports three separate controls:

  1. Acceptance evidence: expected results derived from recorded behaviour rather than a model's interpretation.
  2. A divergence harness: replay infrastructure for parallel comparison.
  3. A coverage map: the observed behaviours and the surfaces that remain unevidenced.

The coverage map is often uncomfortable because production traffic does not exercise year-end close in March. It may miss a leap-year path, a negative balance, or the reinsurance treaty that fires only after a threshold crossed twice since 2009.

Oracle coverage limits every downstream claim. The programme cannot verify behaviour that it never captures. Therefore, the coverage result must appear before an estimate and define what that estimate can mean.

Where coverage is thin, the programme has three legitimate options:

  1. Extend the capture window.
  2. Synthesise inputs for uncovered paths and ask domain experts to adjudicate the outputs.
  3. Accept the gap explicitly and record it.

Discovering the gap in production is not legitimate.

Synthetic cases are challenges, not production observations. Each case records:

  • its generator or source-rule version
  • its author and approval
  • the targeted coverage gap
  • the random seed or input hash
  • the boundary rationale
  • the legacy code and parameter stamps
  • the bound state snapshot
  • the expected-output provenance

The synthetic label survives every aggregate. Generate cases from documented domain constraints and observed schemas or ranges. Obtain the expected result from an isolated legacy execution at the declared state. If execution is impossible, use independent domain adjudication. The migration model may not invent both input semantics and the expected answer. Candidate output may never become its own golden master.

Report observed and synthetic coverage in separate strata. Synthetic cases may challenge boundaries, sentinels, invalid inputs, ordering, retries, and state transitions. However, they do not estimate production incidence or satisfy a requirement for live evidence. Review the suite for over-representation of convenient paths and missing domain classes. A change to the generator, source rule, parameter set, normalisation, or bound state creates a new suite version. It also reopens affected results.

Capture also needs boundaries. The objective is replayable evidence, not an uncontrolled copy of production. The manifest records the source system, legacy release, effective parameter set, extraction window, population basis, completeness, identity mode, and available fidelity dimensions. Sensitive identities can be pseudonymised before leaving the client boundary, but the resulting data remains governed data. Pseudonymisation is a control, not an exemption.

Boundary calls are recorded as effects, not repeated against live partners. A replay can compare the payment instruction that would have been sent without sending it twice. Time, random values, generated identifiers, and unordered output receive explicit normalisation rules whose changes are versioned. A comparison harness that quietly ignores a field has changed the oracle and must say so.

Finally, “the legacy system is the specification” describes its evidentiary role, not its moral authority. Captured behaviour proves what happened under the recorded conditions. It does not prove that the behaviour was intended, lawful, desirable, or complete. That is why §9 separates causal classification from the business's forward-looking decision.

6. The system you are copying does not hold still

Everything in §5 treats the legacy system as a fixed object to be photographed. It is not. A production system remains under change for the entire programme. Lehman's Law of Continuing Change formalised the point: a useful system must keep adapting or become less satisfactory. [20]

The premise of a static copy contradicts the premise of a system worth copying.

6.1 Feature freeze is usually not on offer

Freezing a business-critical system across regulatory deadlines, rate changes, product launches, and defect fixes is rarely available for a multi-year programme. Practitioners call maintaining the old system while modernising it “the big puzzle.” [7]

If an eighteen-month freeze really is available, revisit whether the system matters as much as the business case assumes. Otherwise, plan for a moving target and price it.

6.2 Consequences of a moving target

  1. The oracle decays. A trace is valid for a specific legacy code version and parameter set. A later change can make it stale while it still looks authoritative. Every trace carries both stamps. Affected traces are re-captured. Google practitioners report new features causing false positives in golden-baseline testing and unmanaged expiry of diff suppressions. [21]
  2. Dual maintenance is a meter, not a line item. During parallel operation, assess each business change for both systems and the evidence corpus. Meter rule changes, parameter changes, open backlog, and release freezes separately. No defensible cross-industry benchmark was found. Use the client's change history and expose every assumption.
  3. A third cause of divergence appears. Beyond a candidate defect or a legacy defect, the legacy system may have changed while the target has not yet caught up.

That third class is neither a migration defect nor a legacy bug. Misclassifying it can “fix” the new system back to behaviour the business already retired.

The first consequence is especially dangerous because a stale trace looks stronger than no trace. It arrives with an expected result and fails loudly when the modern system implements the newly approved behaviour. The team then spends effort making current code satisfy obsolete evidence. Version and parameter stamps turn this from a debate into a join: the report can identify which evidence predates which change.

The second consequence belongs in the business case as a rate, not as “parallel running” in a risk paragraph. Use labour hours before converting to a full-time equivalent (FTE) or currency. For month tt, let RtR_t be in-scope rule changes and PtP_t be in-scope parameter releases. Let hR,th_{R,t} include legacy implementation, target re-porting, impact analysis, recapture, comparison, and review. Let hP,th_{P,t} include parameter promotion or synchronization, affected-trace selection, re-evaluation, comparison, and review. A “parameter” change that alters control flow or interpretation belongs in RtR_t, not in the cheaper stream.

The new-work demand is At=Rt×hR,t+Pt×hP,tA_t = R_t \times h_{R,t} + P_t \times h_{P,t}. Add the carrying load λt×Bt1\lambda_t \times B_{t-1} for monthly re-triage, rebase, evidence-freshness review, and coordination. Then convert with the client's productive capacity per FTE-month, HtH_t: FTE demandt=(At+λt×Bt1)/Ht\text{FTE demand}_t = (A_t + \lambda_t \times B_{t-1}) / H_t. Track unfinished assessed work in hours: Bt=max(0,Bt1+AtCt)B_t = \max(0, B_{t-1} + A_t - C_t), where CtC_t is the month's closure capacity after carrying work. During an upstream dependency or target-release freeze, set Ct=0C_t = 0 for the affected scope. The backlog then compounds by new arrivals. Report rule work, parameter work, backlog, carrying load, and frozen months as separate rows. This prevents a low change count from hiding a long queue.

The evidence review found no defensible portable FTE or dollar benchmark for these terms. Derive RtR_t, PtP_t, hR,th_{R,t}, hP,th_{P,t}, CtC_t, HtH_t, and λt\lambda_t from the client's release history, time records, and pilot. Sensitivity-test them rather than replacing missing observations with zero.

6.3 Continuous synchronization without false automation

Every in-scope legacy change creates three linked records before the next scale or cut-over gate:

  1. Legacy change: release, code and parameter versions, effective date, affected boundaries, and approving change record.
  2. Target disposition: ported, not applicable, deliberately deferred, or scope-changing, with an owner and due date.
  3. Evidence disposition: affected traces and calibration items, plus the decision to recapture, re-evaluate, or document why the evidence remains valid.

Continuous integration (CI) can open the target work item, join version stamps, age the backlog, and block a gate. Semantic diffs and traceability maps can nominate affected evidence. They cannot prove that the impact set is complete, so an authorised reviewer owns the final impact decision. At every gate, reconcile the linked records against authoritative release manifests and code, configuration, parameter, schema, and interface-change sources. An unmatched release or unassessed source marks the potentially affected scope STALE. Until the target and evidence dispositions are accepted, that scope cannot cross cut-over.

This is the useful core of a dual-maintenance synchronization pipeline. A diagram that promises automatic semantic completeness would hide the very uncertainty Measurement 6 is supposed to expose.

6.3.1 Cross-language change-port protocol

A legacy and target branch may not share a language, runtime, data model, or transaction boundary. Their changes therefore merge through evidence and intent, not through textual cherry-picking:

  1. Detect and bind. Assign the production legacy change a stable ID. Bind it to the release manifest and source/configuration/parameter/schema/interface hashes. Also bind the effective date, approval, and deployed result.
  2. Build the impact package. Record the changed behaviour and rationale. Record the affected boundaries and data. Include the nominated target components and relevant intent decisions. Also include every trace or calibration item that may now be stale. Unknown impact remains explicit.
  3. Open the target change. Create a linked target pull request (PR) or work item. Give it an accountable owner, due date, impact package, and required evidence refresh. Freeze cut-over for the affected scope. Unrelated work may continue only when dependency analysis supports separation.
  4. Classify each collision. A representation-only adaptation must demonstrate no observable change. Port and recapture a legacy behaviour change. A deliberate target departure consumes the change budget and enters the intent register. An architectural incompatibility enters the §12.1 pivot route. An unresolved conflict is BLOCKED.
  5. Re-prove from fresh state. Rebase the target on the approved business intent, not merely on source syntax. Recapture or re-evaluate the affected legacy evidence. Start both replays from the same declared logical state. Then rerun the applicable coverage, divergence, reconciliation, gate-integrity, and one-way-door checks.
  6. Close with independent acceptance. An independent reviewer may close the linked records only after the evidence resolves. The deployed legacy change, target change, refreshed evidence, comparison result, and review decision must agree. Superseded evidence remains in history rather than being overwritten.

Automation may detect changes, assemble the impact package, nominate dependencies, propose a target patch, run replay, and block a gate. Automated semantic cherry-picking or auto-merge is not evidence of equivalence and may not close the change record. A human accountable owner and independent reviewer remain necessary for the semantic disposition.

6.4 There are two drift rates, and conflating them hides the fast one

What it is How it moves How you detect it
Rule drift The logic itself, a new eligibility test, a changed calculation path Slowly, through the change process, visible in the repository Commit history; the mechanisms in §§6.2–6.3
Parameter drift Rates, thresholds, limits, effective dates, jurisdiction tables Often outside the code-change process Effective-dated source data and owner attestations

A recorded trace is a statement about the logic and the parameters in force when it was recorded. The freshness verdict must report each separately. A single flag averages two different truths into a useful-looking falsehood.

Rule drift usually appears in repositories and change records. Parameter drift often arrives as an effective-dated table owned by finance, operations, or compliance. That difference changes the control. Rule changes trigger impact analysis and recapture. Parameter changes may require re-evaluation rather than full recapture, but only when the trace records which parameter set produced it. “Code current” and “parameters stale” is a legitimate, useful verdict. “Fresh” is not.

6.5 The clock is the real argument for small scope

Drift cost scales with duration. Kulk and Verhoef analyse 84 bancassurance projects and use a volatility of 2% per month. This value is the median of Capers Jones's published industry averages. It is also his recommended default for a project you know nothing about. However, the authors warn that simple failure thresholds were poor predictors. Therefore, the 2% is a starting default, not a measurement of any particular portfolio. Treat it as evidence that movement is material, not as a planning constant. Measure the client's own rate. [22]

Long programmes make the exposure concrete. Commonwealth Bank's CEO described its core-platform programme as more than A$1 billion over five years, then shifted attention to extracting the platform's benefits. [23] A narrow slice reduces the period during which its target can move. The longer analysis sits unused, the more likely it is to become stale.

Do not compound the 2% figure into a forecast without a model of dependence and scope. Requirements are not interchangeable units, and one change can affect many traces. Use the figure only to justify measurement of the local rate. The steering metric compares the migration's closure rate with the combined arrival rate of rule changes, parameter changes, and unresolved recapture work.

6.6 Your instruments drift too

  • Calibration expires. Re-run Measurement 2 when the model, code conventions, or scope changes.
  • Scaffolding expires. Remove a compensating tool when it no longer earns its cost.
  • Experts leave. Availability and succession belong in the control plan, not the staffing footnote.

6.7 What this changes

Drift converts the problem from verification to control. Two systems are moving. The programme must keep the gap bounded, fresh, and attributed.

7. Three questions, six measurements

A process description is insufficient without measurements. Require six, each with a method, denominator, reported value, and stop condition.

The six measurements answer three questions in sequence:

Order Executive question Measurements Gate before proceeding
1 Is it complete? 1. Oracle coverage and 2. Comprehension calibration The scope and denominator are explicit, and every item is accounted for as evidenced, missing, excluded with authority, or UNEVIDENCED. That complete assessment can answer no. Proceed only when no required material item remains missing or UNEVIDENCED, unless the scope is formally reshaped.
2 Is it consistent? 3. Behaviour and service divergence and 4. Data and security posture Every functional, operational, data, and security inconsistency is remediated or recorded as an authorised deliberate departure. Any unexplained critical break or unapproved budget breach blocks the sequence.
3 Where does it say that? 5. Gate integrity and 6. Drift and freshness Every scoped item resolves to protected evidence, a comparison result, an authorised disposition, an owner, and a current date. The resulting record is the deed; it records the canonical decision rather than inventing it.

This is decision order, not implementation delay. Instrument provenance, access protection, sensitivity, and freshness from the first capture. Question 3 assembles and tests the final record after Questions 1 and 2 establish what the record is entitled to say.

For regulated entities, these six measurements answer the inner evidence sequence. They do not by themselves establish compliance. The applicability register in §2 defines the obligation population before Measurement 1, and the obligation-by-obligation assessment uses the traceable evidence after Measurement 6. Neither outer gate collapses into an aggregate equivalence score.

# Measurement Method and reported artifact Stop condition
1 Oracle coverage Record inputs, outputs, boundary effects, workload strata, trust boundaries, and exposed interfaces; map traces to transaction, batch, report, error, and irreversible-effect paths. Report a coverage map plus uncovered surfaces. An irreversible financial or regulatory path, required workload stratum, or material security surface is uncovered with no approved treatment.
2 Comprehension calibration Reconstruct rules blind on modules with known recorded behaviour; report precision, recall, misses, spurious rules, and unscoreable items. Published translation results vary sharply, so the client's result must be measured. [26] COBOL-specific evidence remains preprint evidence. [27] The measured result cannot support the programme's risk tolerance, or the scoring denominator is indefensible.
3 Behaviour and service divergence Shadow matched inputs from the same logical state; retain and classify every candidate difference by field and stratum against a predeclared behavioural budget. Separately measure latency distributions, sustained and burst throughput, concurrency, queueing, resource consumption, saturation, and recovery against a predeclared service envelope. Use legacy self-noise for attribution, never subtraction. Diffy demonstrates the primary/secondary/candidate pattern. [28] Isolate all replay writes: Scientist explicitly warns against experiments that change data. [29] Any unexplained defect-class divergence at cut-over, service-envelope breach, threshold set after results are known, or aggregate adjustment without event-level or workload-stratum attribution.
4 Data and security posture Reconcile records and independently computed business aggregates; review decimal, rounding, overflow, sentinel, date, character-set, and collation semantics. Separately assess the target threat model and exposed interfaces; test authentication, authorisation, injection resistance, secrets and cryptographic handling, dependency risk, audit logging, and resource-exhaustion controls against approved requirements. Any unexplained data break, unresolved critical or high security finding, newly exposed interface without approved controls, or security exception without a named owner and treatment. “Immaterial and unexplained” is two findings, not one.
5 Gate integrity Report provenance separation, post-creation access protection, human separation, and demonstrated sensitivity separately, each with its coverage. Judging evidence lacks independent provenance, changed without explanation, could be relaxed by the migration author, a human-separation claim is unsupported, a gate has not rejected every required seeded-defect class that applies, or a class has not reached the distinct seeded-input count its declared miss-rate bound requires.
6 Drift and freshness Monthly: in-scope rule changes, parameter changes, stale-corpus share, dual-maintenance backlog, and calibration date. Backlog grows across successive periods or evidence exceeds its freshness budget.

How to read the six

Measurements 1–2 answer completeness for the declared scope by reporting both what is evidenced and what is missing. A complete assessment is not necessarily a passing result. Measurements 3–4 test functional, operational, data, and security consistency only after every required material gap is closed or the scope is formally reshaped. Measurements 5–6 make the accepted result traceable and keep it current. The pairing is an executive navigation aid. Every measurement retains its own denominator, artifact, and stop condition.

Measurement 1 covers observable behaviour and declared assurance surfaces rather than source lines. Its map starts with transaction codes, batch jobs, reports, interfaces, error branches, state transitions, outbound effects, workload strata, trust boundaries, and target-only exposed surfaces. A path can execute thousands of covered lines while remaining behaviourally unevidenced if no trace captures the boundary value or resulting write. Report the observed population and named gaps. Full-population language is permitted only when the extract manifest establishes a full population.

For Measurement 2, test the method before using it to assess the system. Choose modules whose behaviour is established, hide that evidence from the extractor, and score reconstructed rules against a human-adjudicated key. Precision without recall rewards a short list that misses rules. Recall without precision rewards invention. Report both measures, enumerate misses and spurious rules, and retain an unscoreable bucket. If semantic matching cannot be automated defensibly, score it manually instead of asking a second model to certify the first.

Published benchmark numbers do not substitute for this calibration. Evaluation design can move the measured number as much as the model does. Researchers corrected one vulnerability-detection benchmark's label noise, duplication, and leakage. They then re-scored models on paired vulnerable and patched functions. A state-of-the-art model's F1 fell from 68.26% to 3.09%. In the most stringent settings, GPT-3.5- and GPT-4-based attempts performed close to random guessing. [45] The study measures vulnerability detection rather than rule extraction. It is cited for the design effect, not for a transferable rate. For the same reason, this calibration is blind and client-run. The programme scores it against an adjudicated key.

Measurement 3 requires a self-noise comparator and isolated writes. Production systems can disagree across matched runs because timestamps, sequences, concurrency, unordered collections, and external state move. Compare legacy-primary, legacy-secondary, and candidate event by event from the same declared logical state. A baseline rate cannot identify the cause of a candidate difference, so never subtract, net, or automatically ignore it. Retain each difference until it is attributed by predeclared field and stratum. Both replay executions capture attempted writes without releasing them. POST, settlement, filing, notification, and other non-idempotent operations belong in the one-way-door register. Report environmental noise, migration defects, legacy defects, accepted changes, and legacy-change lag separately.

The service result is a separate stratum. It is not a footnote on functional agreement. Measure latency as a distribution rather than one average. Under representative workloads, report:

  • sustained and burst throughput
  • concurrent-load behaviour and queueing
  • CPU, memory, and I/O
  • connection or thread-pool pressure
  • saturation and recovery

The legacy envelope provides evidence about current operation. It does not automatically define a target requirement. The charter may demand better or accept worse. However, it fixes each budget, population, environment, and owner before results exist. A functionally identical result outside the approved service envelope does not pass Measurement 3.

Statistical discipline does not require an arbitrary sequential probability ratio test (SPRT). A full declared population is a descriptive census. For a sampled rate gate, the predeclared plan must require affirmative evidence. It can use a one-sided upper confidence bound below the budget. Alternatively, it can use a non-inferiority test whose null is that the rate is intolerable. The plan defines:

  • the sampling unit, populations, and strata
  • acceptable and intolerable rates
  • confidence and error levels
  • power and the sample-size rationale
  • clustering and multiplicity
  • the stopping rule and change authority

A thousand records from one failed batch are not a thousand independent failures. Sequential testing is optional. A pass applies only to the studied population under its assumptions. No rate test overrides the zero-unexplained-divergence gate for critical reconciliation or irreversible-effect paths.

For a reference design with a fixed sample and a one-sided limit, declare four values:

  • the largest acceptable divergence rate p0p_0
  • the smallest intolerable rate p1>p0p_1 > p_0
  • the false-breach probability α\alpha
  • the probability β\beta of missing p1p_1

With δ=p1p0\delta = p_1 - p_0, the normal-approximation starting size is

n0=(z1αp0(1p0)+z1βp1(1p1)δ)2n_0 = \left\lceil \left( \frac{z_{1-\alpha}\sqrt{p_0(1-p_0)} + z_{1-\beta}\sqrt{p_1(1-p_1)}}{\delta} \right)^2 \right\rceil

NIST also gives a continuity correction of 1/δ1/\delta. [39] Use exact binomial calculation or simulation when the approximation is unreliable. Warning conditions include rare events, small expected counts, stratification, or the planned analysis. There is no universal minimum: the client approves p0p_0, p1p_1, α\alpha, and β\beta before results are seen.

Batch or job dependence inflates that starting size. The design effect (DEFF) is the multiplier for this dependence. With mean batch size mm and intraclass correlation ρ\rho, use DEFF=1+(m1)ρ\mathrm{DEFF} = 1 + (m - 1)\rho and plan at least n0×DEFF\lceil n_0 \times \mathrm{DEFF} \rceil record observations. For unequal batch sizes with coefficient of variation CVm\mathrm{CV}_m, a prespecified cluster-level approximation is DEFF=1+((1+CVm2)×m1)ρ\mathrm{DEFF} = 1 + \left(\left(1 + \mathrm{CV}_m^2\right) \times m - 1\right)\rho. [40] The analysis plan states how ρ\rho and CVm\mathrm{CV}_m were estimated and sets a minimum number of independent batches. If those inputs cannot be defended, make the batch the sampling unit or use a prespecified cluster-aware simulation. High transaction volume does not repair a one-batch sample.

If sequential sampling is chosen, apply it to the prespecified independent unit. After nn units and xnx_n divergences, calculate

Sn=xnlnp1p0+(nxn)ln1p11p0S_n = x_n \ln\frac{p_1}{p_0} + (n - x_n) \ln\frac{1 - p_1}{1 - p_0}

Declare a breach when Snln((1β)/α)S_n \geq \ln((1 - \beta)/\alpha), declare the sampled rate within budget when Snln(β/(1α))S_n \leq \ln(\beta/(1 - \alpha)), and otherwise continue. [39] Group-sequential looks operate at the batch level or use a cluster-adjusted design. Predeclare a maximum NmaxN_{\max}. Reaching it without a boundary produces INCONCLUSIVE, unless a separately predeclared terminal test resolves the result.

Measurement 4 reconciles meaning rather than stopping at row counts. In many legacy systems, the data model carries business rules through packed decimal, storage overlays, sentinel values, date windows, character encoding, and collation. Reconcile balances, positions, accruals, counts by state, and other independently computed business aggregates. Never net two breaks into zero: an unexplained positive and an unexplained negative remain two findings. The agent may propose transformations and checks, but the authorised data operator controls execution where corruption would be difficult to reverse.

The legacy system is not a security oracle. Reproducing an old injection flaw is not fidelity. A new application programming interface (API), identity boundary, dependency, or administrative path may have no legacy counterpart to compare. Measurement 4 therefore reports data reconciliation and security posture as separate results. The security result binds a target threat model and attack-surface inventory to abuse cases, automated and manual testing, findings, owners, treatment, and independent review. A scanner's zero findings are no more a security verdict than a green generated test suite is a completeness verdict. The gate must also demonstrate that its security checks reject seeded applicable failures.

Measurement 5 separates four gates. Provenance separation asks whether the migration author could create, select, transform, normalise, or set thresholds for judging evidence. Access protection asks whether that evidence could later be altered, skipped, or relaxed. Human separation asks whether a different person with sufficient authority reviewed the work. Demonstrated sensitivity asks whether the gate can return red at all. Passing one says nothing about the others. Repository identities and review records are screening evidence. They do not prove that a thoughtful review occurred. Report each gate and its coverage separately, such as “provenance separation evidenced for 82%. 18% unevidenced,” rather than the unqualified “independent review passed.”

A broken scoring path can pass the first three gates cleanly. It may have independent provenance, protected thresholds, and separate review while returning a pass for every input, including those it exists to stop. Such a gate can look like a stable suite. Neither the observed rejection rate nor the denominator separates zero actual failures from an eligibility filter that failed to count what it should judge.

Start by reporting each check's eligible count beside its result. An eligible count equal to the full population should be visible for review rather than interpreted automatically as completeness. Next, run an unmutated passing control and seeded-bad inputs through the production scoring path. The minimum sensitivity pack covers every applicable class below. A class marked not applicable needs a recorded rationale and reviewer approval.

  1. Boundary arithmetic: off-by-one thresholds or dates, rounding changes, sign changes, decimal overflow, and precision loss.
  2. Omission and duplication: a dropped record or field, duplicated posting, truncated group, or skipped eligible item.
  3. Representation and order: flipped collation, encoding corruption, reordered causal event, or set-versus-sequence substitution.
  4. Sentinel and error handling: unhandled sentinel, null/default substitution, invalid date, unknown code, or suppressed exception.
  5. State, parameter, and retry: stale parameter set, wrong state snapshot, duplicate retry, lost idempotency key, or transaction-boundary shift.
  6. Authority and side effect: bypassed approval, release of an isolated write, or an unauthorised irreversible effect.
  7. Security and exposure: injection, authorisation bypass, secret disclosure, a newly reachable endpoint, a vulnerable dependency, or resource exhaustion.

Each seeded input records the mutation, expected rejection reason, eligible gate, run ID, result, and most recent rejection date. At least one fixture exercises eligibility and dispatch rather than only the final comparator. The verdict must change for every applicable class, while the unmutated control must still pass. Report the share of gates and required classes without that evidence. For qualitative criteria, separate the fixture author from the criterion author when possible. Otherwise report sensitivity as unevidenced, since both authors may share the same misreading.

The pack also needs a size. One rejected fixture proves a class can fire. However, it says almost nothing about how often that class is missed. For a class with zero misses across nn seeded inputs, the one-sided 95% upper miss-rate bound is 10.051/n1 - 0.05^{1/n}. This bound is near 3/n3/n at any nn worth running. Therefore, ten seeded inputs are consistent with a gate missing three in ten. A hundred bring that to roughly three in a hundred. This is the zero-event case of the exact binomial calculation already required for rare events. It is not a separate method. Before the pack runs, declare the acceptable miss rate for each class. Take nn from that rate and report the achieved bound. Until a class reaches its nn, its result is unresolved rather than sensitive. This applies cross-cutting rule 1 to the gate, rather than the comparison. It preserves the distinction between "not compared" and "no divergence".

nn counts distinct seeded inputs, never repeated executions of one input. The bound assumes nn independent draws. A scoring path that meets the reproducibility requirement returns the same verdict on every rerun. Therefore, repeated executions replay one draw. Even where a model sits in the path, resampling the same input only re-measures that input. It reduces the within-input component of evaluation variance. It provides no evidence about which inputs the gate misses. [46] Sixty executions of six inputs establish the bound at n=6n = 6, not n=60n = 60.

The bound covers only the classes that were seeded. A failure mode that nobody constructed remains outside the bound at every value of nn. Therefore, the programme reports an unevidenced class as unevidenced instead of including it in a favourable total. For the same reason, author separation is part of the control rather than a refinement.

Measurement 6 prevents the first five from expiring. A monthly drift report joins the current legacy release and parameter set to the captured corpus. It lists rule and parameter changes separately. It also calculates stale-corpus share and ages the dual-maintenance backlog. Finally, it records the last comprehension recalibration. One growing backlog period is an alert. Successive growth shows that the target moves away faster than the programme closes the gap.

Measurement 2 can produce useful risk information within days. If the result is poor, narrow the scope, add human extraction, change the method, or stop. Treating calibration as an assumption can commit the programme before the method is tested. A gate that cannot change a decision is ineffective.

Four rules apply across all six:

  1. “Not compared” is not “no divergence.” Missing inputs render UNEVIDENCED.
  2. Every verdict carries its denominator, and every check carries its eligible count. A result over half the changes is a statement about half the changes. A check may report an eligible count equal to the whole population. It may be genuinely universal, or it may miscount what it can judge. The reported rate cannot distinguish these cases. Therefore, report the eligible count beside the result and never infer it from the result.
  3. A proxy stays labelled as a proxy. Author-versus-committer identity can screen for missing review. It does not prove review happened.
  4. One green stratum cannot absorb a red one. Functional agreement does not cancel a service-envelope breach. Data reconciliation does not cancel a security finding. An average across these results is not an admissible verdict.

What makes a measurement admissible

A measurement is admissible only when its artifact can be reproduced and reviewed. Require five properties:

  1. Reproducible: identical inputs produce byte-identical deterministic output.
  2. Derivable: each number traces to a readable rule over stated inputs.
  3. Self-contained: the artifact renders without a network.
  4. Pseudonymisable at source: identities can be salted and hashed before leaving the client boundary. Pseudonymisation reduces exposure. It does not remove the data from data-protection scope. Mixed identity modes are rejected.
  5. No egress exception required: network-dependent collection is a separate operator step that produces a file.

The first two make an artifact evidence rather than assertion. The last three make it producible in the institutions for which it matters.

8. Why per-module success does not mean system success

Per-unit accuracy cannot be carried directly into a system claim.

An IBM research paper illustrates the point with a ten-paragraph COBOL program. At 90% translation accuracy per paragraph, it puts the probability of a fully correct program at about 34%. [30] The multiplication assumes that the errors are independent. The source does not state that assumption. The rest of this section explains why the errors are not independent. The following table extends that illustrative model to units of any kind and computes 0.9ⁿ exactly:

Units in the program P(fully correct) at 90% per unit
10 34.9%
25 7.2%
50 0.52%
100 0.0027%

Inverted, the per-unit accuracy required to reach 95% confidence at the program level:

Units Required per-unit accuracy
50 99.90%
100 99.95%
500 99.99%

The tables are this paper's arithmetic, and they are scenarios rather than forecasts. The source states only the ten-paragraph case. Every other row is computed here. Real errors are correlated and modules differ in criticality. Positive or negative dependence changes the probability materially. Translation-accuracy studies also use different units and cannot be substituted into this table.

The conclusion is narrower and still important: validating each product backlog item (PBI) does not establish system equivalence. Use a system-level oracle and measure critical paths directly. The same IBM paper says formal equivalence was impractical in its context. It also notes that its model judge was more error-prone than subject-matter experts. [30]

8.1 Local success still needs a system-control argument

Component correctness and system control are different claims. Leveson's systems-theoretic safety work explains the distinction. Components can satisfy their local requirements while interactions still permit an unacceptable system state. Hand-offs, timing, feedback, or an incorrect process model can create that state. [47] This paper transfers that mechanism into modernization control design. It does not claim that every modernization is safety-critical. It also does not claim that Leveson's method validates this method. A completed control record does not prove safety, compliance, or outcome.

The source's chain is loss → hazard → safety constraint. The bounded modernization translation is:

  1. Unacceptable outcome. Name the business, customer, financial, operational, or regulatory result the programme must prevent.
  2. Prohibited system state. Name the observable combination of state, authority, timing, and effects from which that outcome could follow.
  3. Enforceable constraint. State what must, may, or must not happen. Name who has stopping authority. Specify what evidence must arrive in time to keep the state controlled.

Every consequential gate should carry the same cross-cutting control record:

  • the outcome, prohibited state, and enforceable constraint.
  • the accountable controller, stopping authority, and required, permitted, and forbidden actions.
  • the controller's process model, assumptions, dependants, and owner.
  • the invalidation or expiry trigger and last-confirmed date.
  • the evidence and feedback channel, maximum useful latency, breach route, and escalation route.
  • hand-off overlaps, gaps, and conflicts, and one end-to-end owner.
  • the exact deployed subject, bound to the executable, configuration, parameters, data contract, environment, and release.

Review the gate against four ways a control action can fail:

  • the required action is omitted
  • an action occurs when it is unsafe
  • the timing or sequence is wrong
  • the action stops too soon or persists too long

Those categories come from Leveson's control analysis. [47] Their modernization use is this paper's adaptation, not a regulator-authored taxonomy.

Approval is therefore a maintained state, not a ceremony. Any changed artifact, configuration, assumption, authority, or operating condition reopens the affected constraint. So do an anomaly, a discovered false green, a waiver, or a near miss. The named owner must revalidate the constraint or block the dependent decision. Closure requires evidence that the corrective action was effective. A historical green result tied to a different deployed subject cannot authorise the current one.

The remaining problem is a business decision rather than a measurement problem.

9. When the two systems disagree, the old one is not automatically right

Everything up to here has been built to detect divergence. This section is about what to do with one, and it is where a conversion is most often quietly lost.

The oracle establishes what the system did under captured conditions. It has no authority to decide what the system should do next.

The legacy system is the best available evidence about the past. It is not a specification for the future. Those claims are different. A method that conflates them will faithfully reproduce decisions made by people who left in 2004 for reasons that expired in 2011. It will do so at considerable expense.

9.1 The triage has two questions, and the second one is the authoritative one

Question 1, for engineering: Why do the systems differ?

Cause
Defect in the new system The translation is wrong
Faithful reproduction of a legacy bug The translation is right; the source was wrong
Legacy moved The old system changed and the new one has not caught up (§6.2, point 3)
Forced platform divergence Rounding, collation, or ordering behaviour that cannot be reproduced (§1.3)
Non-determinism A candidate difference attributable to a predeclared non-deterministic field and matched legacy-secondary behaviour; the aggregate baseline alone is insufficient

Question 2, for the business: Which behaviour should govern future operation?

No artifact can answer the second question. Code, traces, and models contain past behaviour rather than current business intent. A person with authority must supply that intent.

The business answers Disposition
The old behaviour is right Defect. Fix the new system.
The new behaviour is right An improvement. Accept the divergence. Record it as a deliberate behaviour change with a named owner.
Each is right, for a different consumer A boundary split. Preserve the old behaviour where it is contractual and adopt the new one elsewhere. Every side carries its own named owner, and the split itself is the deliberate change.
Neither is right You have found a requirement, not a bug. Route it out of the migration entirely.
Nobody can say Escalate. Do not default.

The answer may be unavailable. Practitioner interviews cited in §6 describe the loss of business knowledge directly:

"People don't know the rules anymore because they never use them because the systems do the work. So there is no business knowledge anymore in the business." "Nobodies know all the rules anymore which are in the system."

[7]

The problem compounds over time. Documentation may be incomplete, and code records what happens without preserving the original rationale. After decades of automation, the people who once knew the rule may no longer need to remember it.

Present intent may therefore need to be constructed from the evidence surfaced by the migration. That work requires current decision-makers and belongs in the delivery plan as scheduled business activity, not as an engineering exception queue.

9.2 The last row is the one that costs money

"When in doubt, match the legacy" appears conservative, but it still makes a decision. An engineer under delivery pressure may preserve behaviour that nobody could justify because a 4 p.m. divergence was never escalated.

Some legacy behaviour remains operationally essential. Other behaviour may be a workaround for a retired system. It may also be a rounding convention that predates the euro. Alternatively, it may be a poorly automated step from 1997 that three departments now use in their plans. Preserving such behaviour carries an implementation cost and an ongoing operating cost.

Migration creates an unusually cheap inspection point for some changes because the behaviour is already under examination. It is not automatically the cheapest implementation point: extra scope, validation, and schedule risk can make deferral rational. Refusing to decide is still a decision to carry the debt forward.

9.3 The boundary that keeps scope controlled

Uncontrolled improvements can consume the migration schedule. A fixed boundary makes deliberate change possible without turning the conversion into an unplanned rewrite:

Fidelity remains the default, but the business may rebut it through a controlled decision:

  1. Fidelity is the default, because most behaviour has never been decided about by anyone currently employed. Reproducing it is the correct treatment of behaviour nobody has an opinion on.
  2. A deliberate behaviour change requires a named owner and a written rationale. Not a ticket comment. A named human who will answer for it at cut-over.
  3. Changes consume a declared change budget, fixed before the pilot starts and expressed as a count. Once the budget is exhausted, further ideas move to the backlog.
  4. Everything above the budget goes to the post-cut-over backlog. That backlog is the exercise plan for the option described in §1.2. The change budget keeps improvements from becoming unplanned migration scope.

9.4 The intent register

The comparison report records every difference. Its defect/noise budget governs accidental divergence. A separate change budget governs deliberate differences. A freshness budget governs legacy-change lag. Deliberate differences also need a decision record:

The intent register records every behaviour deliberately not preserved. Each entry names the owner, rationale, date, evidence that made the prior behaviour wrong, and disposition selected from §9.1.

After the completeness gate has closed and every inconsistency has been resolved, these entries become part of the canonical decision record. They do not make the legacy output true or turn a preference into a discovered fact. They record which behaviour an authorised person selected for the future, why, and against which evidence—the deed that answers “Where does it say that?”

The evidence field is causal. A list of documents considered does not reveal what changed the decision-maker's mind. A reviewer in year three needs the evidence that made the prior behaviour wrong. If nothing did and the business simply preferred the new output, record that fact. It is more useful than a citation that played no part in the decision.

Record the disposition from a fixed set rather than as prose. An entry such as "agreed with business, moving on" carries no reproducible decision. The fixed set distinguishes an accepted improvement from a routed-out requirement. The improvement stays in the migration and consumes change budget. The requirement leaves the migration and consumes none. Those decisions have different costs and need different records.

It is a living document, not a closeout artifact. Three groups use it:

  • the business, to see what it is buying
  • the auditor, to see that people decided the changes rather than drifted into them
  • the team in year three, to determine why the new system does not match a thirty-year-old report

Without the register, deliberate improvements and undetected defects both appear as "the systems differ." A conversion can then become a re-implementation without an explicit decision.

The register has a limit. At the time of entry, a genuine requirement change may look identical to an exception granted after an unwelcome result. Either can have an owner, date, and citation. The difference appears in subsequent treatment rather than in a single field. The register preserves accountability, but it cannot establish causality on its own.

Three controls recover most of the distinction. The disposition sends a requirement into separately justified work while an exception remains in the migration. The change budget, fixed in the charter before results exist, prices the exception. The routed-out register keeps removed work visible so a reviewer can detect when an exception has been reclassified to avoid the budget.

This does not make goalpost movement impossible. It assigns a cost that the charter priced before anyone knew where the goalposts would be. It also makes that cost visible to someone who was not in the room. This claim is weaker than "the register prevents it," and it is the true one.

9.5 What this does to "behavioural equivalence"

Behavioural equivalence becomes a risk-management default rather than the programme goal. It is the appropriate treatment for behaviour that nobody has reconsidered. It does not establish that rules written in 1994 remain correct.

A method that treats equivalence as the objective will resist every improvement the business asks for and call that rigor. A method with no default will accept every one of them and call that responsiveness. The discipline is a rebuttable default with a budget and a register.

10. Legacy experts as a control

IBM Research, reporting on its own COBOL-to-Java work, states that translated code cannot simply be trusted and that manual validation remains necessary. [31] Some required decisions have no answer in the code.

A replay harness can find a difference. It cannot decide whether it is:

  • a rule the business depends on, which must be preserved
  • a previously unnoticed bug that the business may now choose to fix, with an associated cost
  • a bug the business has silently worked around for years, where fixing it breaks the workaround
  • a behaviour that was intentional in 1994 for a reason that no longer exists

Bug-for-bug compatibility is a decision, adjudicated per rule. Past intent was never compiled. Present intent has not been written down yet.

The adjudication needs someone who knows what the system was built to do and someone authorised to say what it should do now. A programme staffed only with the first preserves everything. One staffed only with the second can rewrite the business by accident.

Where an applicable regime requires independent review, the reviewer must also have the standing to challenge development. A reviewer who cannot force a change is not a gate.

That authority has a cultural cost. A reviewer who stops a multi-million-dollar cut-over over one unresolved semantic difference will create schedule pressure, political friction, and demands to reinterpret the threshold. The sponsor must pre-commit—before results exist—to the escalation path, decision time, reviewer protection, and schedule and funding reserve for a negative verdict. A gate that is binding on paper but punishes the person who uses it is not an operating control.

Name the domain experts and business decision owner. Budget their time through cut-over and one full business cycle, including annual paths. Put succession and retention in the control plan because their time supplies the adjudication mechanism for §9.

10.1 Make human review an operating control

“A human reviews it” is not yet a control. Human oversight can suffer automation bias, misplaced confidence, and confirmation effects. A safety-case synthesis and peer-reviewed secure-coding study document two versions of the problem. [32] [33] A controlled clinician study documents a third. [34]

The domain transfer is an inference: a pathology experiment does not measure modernization review. Its value is the mechanism it exposes. A reviewer shown a polished rule inventory may accept an error that agrees with an existing assumption because the presentation feels coherent. The remedy is not “more human in the loop.” It is a review design that makes disagreement discoverable.

Make review measurable:

  • Adversarial framing. The question is not "does this look right?" but "find the rule that is missing." The first question has a comfortable answer.
  • Held-out cases. Present traces from Measurement 1 that the reviewer has not seen and the extraction did not use. Ask what the system does. Compare to what it did.
  • A measured catch rate. Seed known discrepancies into review batches and record how many the reviewer finds. A review process with no measured catch rate has an unknown value. Unknown values do not belong in a safety case.

The responsible, accountable, consulted, and informed (RACI) responsibility matrix in Appendix B makes authority visible. It does not prove that the control operated. Evidence of operation is separate. It includes attendance or review records, decisions, reviewer-required changes, unresolved challenges, and the measured catch rate. The migration author is not accountable for validation. A named human with stopping authority is accountable for an irreversible write.

A measured catch rate, even a mediocre one, provides more control than an oversight step whose effectiveness has never been tested.

11. The door that only opens one way

Version control makes code rollback routine through a tag, revert, and redeployment. State and external effects require different treatment. The modern system may post a transaction or allocate an identifier from a new sequence. It may send a settlement instruction or emit an event on which a downstream consumer acts. In each case, code reversion leaves the new state in place. The result may be worse than either system alone.

Any migration therefore contains moments that version control does not protect. Every non-idempotent internal or external effect requires a person-linked human authoriser before release. The first irreversible external write is a release-specific hard gate that is never delegated.

Before cut-over, classify every outbound effect in writing:

Effect class Treatment
Idempotent, internally scoped Roll back with the code
Non-idempotent internal (sequences, IDs, accruals) Named human authoriser. Dual-write with reconciliation. Documented and tested compensation.
External and reversible (some payment rails, internal messaging) Named human authoriser. Test the compensating transaction before cut-over. Do not design it during the incident.
External and irreversible (settlement, regulatory filing, customer notification) Hard gate. Human authorisation. No agent authority, at any maturity level.

"We can roll back" is only true of the first row. A rollback story that rests on Git describes the code and says nothing about the ledger.

The register makes the claim testable. For every outbound effect, it records:

  • the boundary and effect class
  • the reversal or compensation
  • evidence that the compensation was tested
  • the owner and person-linked human authoriser
  • the gate status

A non-idempotent internal or external effect without that authoriser is BLOCKED. A non-idempotent internal or external-reversible effect without tested compensation is UNEVIDENCED. It cannot pass its gate. An external-irreversible effect remains pending until the human authoriser approves the specific release. The tool may validate those fields. It cannot discover the business effect or invent the compensating transaction.

Cut-over planning then separates four questions that are often collapsed:

  1. Can the code be redeployed or reverted?
  2. Can state written inside the new boundary be reconciled or compensated?
  3. Can an external party's action be reversed after it has observed the write?
  4. If rollback or compensation fails or is impossible, what tested fix-forward path restores safe operation, state integrity, and reconciliation—and who can authorise it?

Deployment tooling answers only the first question. The programme places the gate before the earliest “no” in the remaining questions. The human authoriser records approval for the specific release, evidence window, and effects. The programme does not receive permanent approval.

Fix-forward is not permission to design a repair during the incident. Before cut-over, rehearse:

  • the known-good deployment path and traffic isolation
  • state or schema repair and reconciliation
  • evidence refresh and decision rights
  • the maximum recovery time for the release

An irreversible effect may have no safe rollback or compensation. In that case, the human authoriser cannot open the gate without the tested fix-forward package and release-specific risk acceptance.

12. What the pilot is for

A useful pilot is designed to expose failure cheaply. A demonstration arranged to succeed shows only that a favourable case can be staged. The pilot instead tests whether the method survives a bounded case before the programme places the business at risk.

Publish stop and scale criteria before the pilot begins. The criteria below provide a framework rather than universal numeric thresholds. Before evidence is observed, the signed charter records each client value, its source or calculation, accountable owner, approval date, and change-control rule. It also fixes scope, duration, funding, risk appetite, data authority, regulatory applicability, and the split between shipped tooling and manual controls.

Default criteria to negotiate against that risk tolerance:

Stop the current path if:

  • Oracle coverage cannot be brought above the agreed floor for the paths with irreversible effect
  • Comprehension calibration misses the predeclared recall floor for a critical behavioural surface, or system-level oracle results remain below the approved floor. Recall is not compounded across modules as if it were an independent correctness probability
  • Divergence rate in the defect class does not converge over successive replay windows
  • Any data reconciliation break remains unexplained
  • A required service-envelope budget is breached under its declared workload, or capacity and recovery evidence is missing
  • A critical or high security finding remains unresolved, or a new exposed surface lacks approved controls
  • The acceptance corpus shows unexplained modification
  • The dual-maintenance backlog grows across successive reporting periods. The new system is losing the race against the old one, so scope must narrow or the programme must end (Measurement 6).
  • The change budget is exhausted during the pilot. If a single module's divergences consume the allowance for deliberate behaviour changes, the programme is not a conversion. That finding concerns scope, not execution failure. It means §1.1 was answered incorrectly. Re-decide the programme shape rather than raise the budget.

Proceed only when all six measurements report visible denominators, eligible counts, and gaps. Each must be within its client-approved budget. No relied-on gate may be sensitivity-unevidenced. Classify and accept remaining divergences by name. Every deliberate behaviour change needs an owner in the intent register. Service and security results must pass as separate strata. Data and regulatory approvals must be current. Each non-idempotent or external effect needs a person-linked human authoriser. Each effect claimed reversible needs tested compensation. The release also needs a rehearsed fix-forward package for rollback or compensation failure.

There is a third outcome: this system, at this readiness, should not be modernized this way. It is a legitimate pilot result, not a failure. The substrate may not support the method. Money may be better spent on encapsulation, targeted replacement of a bounded capability, or retention for another two years. A method unable to return that answer is a sales process with stages.

12.1 Stop is not one outcome: pre-agree the off-ramps

A breached gate should not trigger improvised abandonment or an automatic increase in tolerance. The charter assigns each material breach to one of three authorised outcomes before sunk cost begins arguing for continuation:

Outcome Typical trigger Treatment of remaining scope and budget
Reshape A bounded, non-critical surface lacks coverage or calibration, while the critical core remains evidencable Narrow the automated-conversion boundary; move the excluded surface to manual porting, later replacement, or explicit retention; rebaseline every affected budget and benefit
Pivot Direct conversion does not converge, or platform semantics/coupling make it the wrong architecture Stop translating the affected capability; evaluate encapsulation, a strangler slice, re-implementation, or package sourcing; require a new business case and gates rather than carrying conversion approval across
Stop / retain A critical reconciliation break remains unexplained, evidence integrity is compromised, required authority is absent, or the investment case no longer survives End the conversion path; preserve or delete captured evidence under the approved retention plan; document the safe operating, replacement, or retirement decision

These are decision routes, not euphemisms for passing a failed pilot. The original breach remains in the record. A reshape requires a versioned new charter whose scope and budgets are frozen prospectively. It then needs validation on a fresh holdout or replay window. Manual ports retain their own oracle, reconciliation, and one-way-door gates. A pivot is a new programme shape under §1.1. Stop/retain still needs an owner, risk treatment, and review date.

12.1.1 Contract the right to stop or pivot

An off-ramp is not executable if the statement of work assumes conversion volume. It also fails if the provider withholds evidence needed to continue elsewhere. The same applies if every pivot becomes an unpriced negotiation. The client's procurement and legal functions must adapt the following minimum commercial controls. The pilot term sheet should contain them:

Contract control Minimum provision
Funding tranches Cap pilot spend and make each later tranche contingent on named evidence gates and a fresh approval. Do not guarantee estate-wide conversion volume before the pilot result.
Evidence and harness custody Give the client custody of and sufficient perpetual, transferable rights to continue using client-funded corpora, manifests, mappings, adapters, schemas, normalisation rules, budgets, thresholds, fixtures, reports, and decision records. Supply the replay, divergence, service, and security harnesses—or source, build instructions, dependency pins, runbooks, and interfaces sufficient to reproduce their outputs—in documented exportable formats inside the client's controlled environment. A provider-hosted black box cannot be the sole evidence path. Data-protection, confidentiality, source-licensing, supplier pre-existing IP, and third-party restrictions are listed separately before approval.
Pivot and residual-budget rights Permit unused committed funds to be redirected to an approved reshape, encapsulation, re-implementation, package evaluation, or orderly retention plan through a new charter. A pivot does not carry the failed conversion acceptance forward.
Termination assistance Pre-price and time-box evidence export, configuration and dependency handover, knowledge transfer, credential/access revocation, subcontractor exit, and return/deletion attestations required by the retention plan.
Stranded-cost disclosure Identify non-cancellable licences, environments, minimum commitments, data-egress charges, and exit fees at each decision gate, with an owner and maximum exposure.
Acceptance integrity A stop or reshape is not deemed technical acceptance. The original breach remains visible. Thresholds cannot be relaxed retroactively. A commercial dispute cannot authorise a blocked live or irreversible effect.

The approval paper should show maximum committed spend only through the next decision gate. It must also show each off-ramp's cost and who can exercise it. That turns “stop” from a methodological aspiration into a funded and contractually available board decision.

12.2 Estimate timelines only after a probe

Portable speedup percentages describe someone else's codebase, against a baseline that is rarely defined. There is a sharper reason to distrust them:

METR ran a randomized controlled trial with 16 experienced open-source developers and 246 tasks in mature repositories. [35]

Estimate
Developers, before starting 24% faster
Developers, after finishing the tasks 20% faster
Economics and ML expert forecasters 38–39% faster
Measured outcome 19% slower

Every group got the sign wrong, including the developers reporting on work they had just personally completed.

The contrast case matters: a controlled experiment on GitHub Copilot measured 55.8% faster completion on a specified, greenfield JavaScript HTTP-server task. [36] The studies do not conflict. They studied different tasks and populations.

Legacy modernization resembles the second shape: work in a mature repository with hidden constraints, unfamiliar interactions, and costly validation. The cited studies do not establish that it is the most extreme case, and greenfield-task speedups cannot simply be transferred to it.

At the delivery level, DORA's 2024 survey associates each 25% increase in AI adoption with two estimated changes. Throughput decreases by 1.5%, and delivery stability decreases by 7.2%. The study uses observational survey modelling, not an RCT. It does not establish that AI caused either change. [37]

Average risk is insufficient. Flyvbjerg and Budzier reported a 27% mean cost overrun across 1,471 IT projects, with one in six showing much larger tail outcomes. [38] A point estimate hides the case most likely to threaten the business.

Run Measurements 1 and 2 on a bounded module before estimating. Use client data and present a distribution with its tail rather than a single figure. Some work cannot be estimated responsibly until it has been probed. Identifying that limit in week one is cheaper than discovering it in month nine.

13. Claim boundaries

The method has eight non-negotiable limits. The earlier sections contain the proof.

  • It reports measured divergence against a predeclared budget, not behavioural identity or zero risk.
  • Test coverage is a screening metric, not proof of correctness. Demonstrate fault-detection effectiveness separately. [24] [25]
  • Calibration reports what the method recovered on known modules. It does not claim that an AI “understood” the estate.
  • Timelines and speedups are estimated only after probing the client system (§12.2).
  • Validation is ongoing monitoring, not a certificate, where the applicable regime requires outcomes analysis. [11]
  • Regulatory requirements and this paper's inferences stay separate. Human-review independence may be binding in scope. Machine separation is a proposed control, not a quoted rule.
  • Traceability is not compliance. Any compliance conclusion is obligation-by-obligation, owned by qualified human judgment, and limited to a named entity, role, system, jurisdiction, instrument set, obligation population, and date.
  • Functional and data equivalence do not establish service quality or security. Operational envelopes and target security posture keep their own populations, budgets, findings, and stop conditions.

Unknown coverage, reviewer identity, freshness, or artifact provenance renders UNEVIDENCED. It never silently becomes zero, complete, or passed.

14. What this costs you in honesty

The method imposes four costs that belong in the decision.

Visible progress starts later. Capture, calibration, and reconciliation design consume weeks before a line of modern code exists. Translation can begin sooner, but it then proceeds against an unmeasured target.

Incremental migration adds complexity. Bisbal and colleagues rejected big-bang cut-over as unrealistic because of risk, while warning that incremental migration's inherent complexity might increase overall migration risk. [16] Parallel schemas, dual writes, adapter layers, and a long dual-maintenance window are the price of avoiding a single cut-over bet.

The trade may still be correct, but it is not free. No large-N empirical comparison of the two approaches supports a universal outcome claim. The preference for incremental migration rests on risk reasoning rather than comparative outcome data.

Business decisions require recurring attention. A change budget, intent register, and divergence triage require an authorised business decision-maker at the cadence differences arrive, often weekly rather than quarterly. The provider cannot absorb that responsibility without making silent business decisions on the client's behalf.

Bad news arrives on a schedule. Measurement 6 may show in month four that the target is losing the race against the legacy system. The gate must remain intact even when the result is uncomfortable, because the value of stopping depends on stopping early.

15. Decision synthesis

Approve the next commitment only when it changes a decision rather than merely reporting activity. The decision test asks whether the accountable executive can identify:

  • the declared population and its remaining gaps
  • the separate functional, service, data, and security results
  • the named person who accepts each deliberate departure

If not, the evidence may guide more work but cannot authorise scale or live effects.

For a proposed irreversible write, apply a second test. The release owner must identify the exact effect, its human authoriser, tested compensation or fix-forward path, and the evidence bound to that release. If any answer is missing, the effect remains blocked. For a regulated entity, the qualified decision-maker must also identify the applicable obligation and its disposition before making a positive compliance claim.

This is a decision discipline, not a conformity assessment. It keeps the claim bounded to the stated population, evidence, date, and authority. That boundary makes a negative finding actionable: reshape, pivot, retain, or stop before a dashboard turns an unresolved gap into a release decision.

Pierre Boutquin · Version 1.8 · 3 September 2026 · Public

Appendix A — Numbered evidence register

† Preprint, not peer reviewed. A dagger on a claim in the body means the supporting source is a preprint: submitted but not yet through peer review. It is cited because it is the best available evidence on the point, and it is marked because its evidential weight is lower than a peer-reviewed or official source. Each entry below repeats the status in full.

Citation rule. Body citations identify the claim they support. Every entry below records source type, publication date, version, pinpoint, access date, and the permitted use. A preprint label is repeated at first use in the body and here. Unless stated otherwise, links were accessed 15 August 2026.

[1] IBM, z/OS 3.1: Planning for Sub-Capacity Pricing. Type: official vendor manual · Published: updated 17 May 2024 · Version: SA23-2301-60 · Pinpoint: “Traditional sub-capacity pricing,” definition of highest observed four-hour rolling average · Accessed: 15 August 2026 · Supports: R4HA basis of traditional sub-capacity MLC. PDF

[2] Martin Fowler, “Original Strangler Fig Application.” Type: first-party practitioner essay · Published: 29 June 2004 · Version: web edition, undated revisions · Pinpoint: paragraphs describing new functionality delivered before migration completes and value retained if work stops partway · Accessed: 25 August 2026 · Supports: incremental delivery can create value before a migration finishes; transitional-architecture estimating is this paper's inference. Source

[3] IBM, Enterprise COBOL for z/OS 6.4 Language Reference. Type: official language specification · Published: 2024 · Version: 4th ed., SC27-8713-03 · Pinpoint: REDEFINES p.225; PACKED-DECIMAL pp.238–239; ROUNDED p.296; EBCDIC Appendix C · Accessed: 25 August 2026 · Supports: decimal, rounding, storage-overlay, and collation semantics. PDF

[4] IBM Redbooks, Modernizing Applications with IBM CICS. Type: official vendor technical paper · Published: 2020 · Version: REDP-5628-00 · Pinpoint: chapter 2, pseudo-conversational programming and COMMAREA · Accessed: 15 August 2026 · Supports: CICS task and state semantics. PDF

[5] IBM Redbooks, VSAM Demystified. Type: official vendor technical manual · Published: 2022 · Version: 3rd ed., SG24-6105-02 · Pinpoint: chapter 1 §§1.3–1.5, especially pp.5–8, 19–20 and 29–33 · Accessed: 25 August 2026 · Supports: application behaviour coupled to key, relative-record, relative-byte, and physical-position access. PDF

[6] AWS Prescriptive Guidance, Guide for AWS Large Migrations, “About the migration strategies.” Type: official vendor guidance · Published: 28 February 2022; last revised 2 May 2022 · Version: PDF edition; the live web edition agrees verbatim · Pinpoint: “About the migration strategies” overview, p. 14; “Refactor or re-architect” final paragraph, p. 19 · Accessed: 25 August 2026 · Supports: refactoring is not the default for large migrations. Source

[7] Batlajery et al., Industrial Perception of Legacy Software Systems and Their Modernization. Type: university technical report; interview and survey study · Published: 2014 · Version: UU-CS-2014-004 · Pinpoint: practitioner quotations in results/discussion on concurrent maintenance, knowledge loss, and feature freeze · Accessed: 15 August 2026 · Supports: moving-target and dual-maintenance mechanisms. PDF

[8] OSFI, Guideline E-23: Model Risk Management (2027). Type: official supervisory guideline · Published: 11 September 2025; effective 1 May 2027 · Version: final E-23 · Pinpoint: A.4 “Model”; Principle 3.4 “Model review” · Accessed: 15 August 2026 · Supports: scope, modification triggers, and independent review for in-scope FRFI models. Source

[9] PRA, Model Risk Management Principles for Banks. Type: official supervisory statement · Published: May 2023; effective 17 May 2024 · Version: SS1/23, issued with PS6/23 · Pinpoint: scope ¶¶1.2–1.5; Principle 4.1(d)–(e) · Accessed: 15 August 2026 · Supports: validation independence and organisational standing in the statement's scope. PDF

[10] FINMA, Governance and Risk Management When Using Artificial Intelligence. Type: official supervisory communication; observed findings, not rule text · Published: 18 December 2024 · Version: Guidance 08/2024 · Pinpoint: §2.7, p.7 · Accessed: 15 August 2026 · Supports: observed gaps between AI development and independent review. PDF

[11] Federal Reserve, FDIC & OCC, Revised Guidance on Model Risk Management. Type: official interagency supervisory guidance; non-binding · Published: 17 April 2026 · Version: SR 26-2; supersedes SR 11-7 and SR 21-8 · Pinpoint: applicability statement; §§II–III and VI; footnote 3 · Accessed: 15 August 2026 · Supports: covered banking populations, model and AI scope exclusions, effective challenge, independence, and ongoing outcomes analysis. Landing page · PDF

[12] European Central Bank, ECB Guide to Internal Models. Type: official supervisory guide · Published: 19 February 2024 · Version: release 3.1 · Pinpoint: general topics §§15 and 19, pp.28–29 · Accessed: 15 August 2026 · Supports: separate validation and development staff for internal models in scope. PDF

[13] Regulation (EU) 2024/1689, Artificial Intelligence Act, as amended by Regulation (EU) 2026/1744. Type: primary legislation · Published: original 12 July 2024; amendment 24 July 2026 · Version: Official Journal texts, CELEX 32024R1689 and 32026R1744, in force at 15 August 2026 · Pinpoint: Articles 6, 17 and 113; Annexes I, III and IV §2(a); amending Regulation Article 1(40) · Accessed: 15 August 2026 · Supports: high-risk scope, provider documentation duties, staged application dates, and documentation of third-party development tools; does not state an AI-authored percentage. AI Act · 2026 amendment

[14] PREPRINT: NOT PEER REVIEWED. Pavuluri et al., “ScarfBench.” Type: research preprint · Published: submitted 7 May 2026; revised 18 May 2026 · Version: arXiv:2605.06754v2 · Pinpoint: Abstract; §5.1; Tables 3–4 · Accessed: 15 August 2026 · Supports: benchmark design, 204 tasks, pass rates, and one fully equivalent target. Source

[15] PREPRINT: NOT PEER REVIEWED. Liu et al., “MigrationBench.” Type: research preprint · Published: submitted 14 May 2025; revised 28 May 2026 · Version: arXiv:2505.09569v3 · Pinpoint: Abstract; evaluation results for the selected 300-repository subset · Accessed: 15 August 2026 · Supports: Java 8→17 minimal and maximal pass@1. Source

[16] Bisbal et al., “Legacy Information Systems: Issues and Directions.” Type: peer-reviewed journal article · Published: September/October 1999 · Version: IEEE Software 16(5) · Pinpoint: p.107, cut-over and incremental-migration discussion · Accessed: 25 August 2026 · Supports: incremental migration reduces cut-over exposure while adding coordination and coexistence complexity; it is not cited for the issue editor's separate rediscovery wording. DOI

[17] Comella-Dorda et al., A Survey of Legacy System Modernization Approaches. Type: government-funded technical note · Published: April 2000 · Version: CMU/SEI-2000-TN-003 · Pinpoint: pp.3–6, legacy-system value and replacement risk · Accessed: 15 August 2026 · Supports: encapsulated business expertise and robustness uncertainty. PDF

[18] PREPRINT: NOT PEER REVIEWED. Ahmed & Galib, “AgentModernize.” Type: research preprint on eight synthetic scenarios · Published: submitted 17 May 2026; revised 4 August 2026 · Version: arXiv:2605.17535v2 · Pinpoint: Abstract; Tables 3–4; limitations in §VI · Accessed: 15 August 2026 · Supports: 23.0% mean BER for one configuration and the limits of the result. Source

[19] Barr et al., “The Oracle Problem in Software Testing: A Survey.” Type: peer-reviewed survey · Published: May 2015 · Version: IEEE Transactions on Software Engineering 41(5) · Pinpoint: pseudo-oracles subsection, pp.507–525 · Accessed: 15 August 2026 · Supports: independently produced alternative implementations as pseudo-oracles. PDF

[20] M. M. Lehman, “Laws of Software Evolution Revisited.” Type: peer-reviewed conference chapter · Published: 1996 · Version: LNCS 1149 · Pinpoint: Law I and Law II, pp.108–124 · Accessed: 15 August 2026 · Supports: continuing change and increasing complexity. PDF

[21] Gulzar, Zhu & Han, “Perception and Practices of Differential Testing.” Type: peer-reviewed empirical study · Published: May 2019 · Version: ICSE-SEIP 2019 · Pinpoint: results on false positives from new features and expiry of diff ignores, pp.71–80 · Accessed: 15 August 2026 · Supports: golden-baseline decay and suppression-management gaps. PDF

[22] Kulk & Verhoef, “Quantifying Requirements Volatility Effects.” Type: peer-reviewed journal article · Published: July 2008 · Version: Science of Computer Programming 72(3) · Pinpoint: empirical results and threshold analysis, pp.153–161 · Accessed: 15 August 2026 · Supports: measured volatility and the warning against deterministic thresholds. PDF

[23] Brett Winterford, “CommBank Declares Core Bank Overhaul Complete.” Type: trade-press report quoting a primary AGM speech · Published: 30 October 2012 · Version: exact article permalink · Pinpoint: paragraphs 1–9, quoting CEO Ian Narev · Accessed: 15 August 2026 · Supports: more than A$1 billion, five years, and post-platform benefits focus. Article

[24] Inozemtseva & Holmes, “Coverage Is Not Strongly Correlated with Test Suite Effectiveness.” Type: peer-reviewed empirical study; ACM Distinguished Paper · Published: June 2014 · Version: ICSE 2014 · Pinpoint: methodology and size-controlled results, pp.435–445 · Accessed: 15 August 2026 · Supports: coverage is an imperfect proxy for fault detection. PDF

[25] PREPRINT: NOT PEER REVIEWED. Zhao, Zhou & Cohen, “Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness?” Type: replication-study preprint · Published: 24 July 2026 · Version: arXiv:2607.22880v1 · Pinpoint: Abstract; results comparing regression-style and buggy-code settings · Accessed: 15 August 2026 · Supports: context-dependent limits of coverage and mutation proxies. Source

[26] Pan et al., “Lost in Translation.” Type: peer-reviewed empirical study · Published: ICSE 2024; preprint revised 16 January 2024 · Version: arXiv:2308.03109v3 / DOI 10.1145/3597503.3639226 · Pinpoint: Abstract; results over 1,700 samples · Accessed: 15 August 2026 · Supports: 2.1%–47.3% correct translations in the studied language pairs; not COBOL-specific. Source

[27] PREPRINT: NOT PEER REVIEWED. Dau et al., “COBOL-Coder.” Type: research preprint · Published: 5 April 2026 · Version: arXiv:2604.03986v1 · Pinpoint: Abstract; COBOLEval results · Accessed: 15 August 2026 · Supports: GPT-4o 41.8% compilation success and 16.4 Pass-1 in that benchmark. Source

[28] Twitter, Diffy. Type: open-source software documentation · Published: repository created 2015; archived 23 February 2021 · Version: archived README on master · Pinpoint: noise-cancellation design and ignored mutating HTTP methods · Accessed: 15 August 2026 · Supports: primary/secondary/candidate baseline pattern. Repository

[29] GitHub, Scientist. Type: open-source software documentation · Published: repository created 2014 · Version: README on main, live at access date · Pinpoint: “How do I science?” safety warning · Accessed: 15 August 2026 · Supports: experiments that change data are unsafe without isolation. Repository

[30] PREPRINT: NOT PEER REVIEWED. Froimovich et al., “Quality Evaluation of COBOL to Java Code Transformation.” Type: IBM research preprint · Published: 31 July 2025 · Version: arXiv:2507.23356v1; submitted to ASE 2025 · Pinpoint: composition example and evaluation limitations · Accessed: 15 August 2026 · Supports: illustrative independent-error arithmetic, impracticality claim in context, and model-judge limitation. Source

[31] PREPRINT: NOT PEER REVIEWED. Hans et al., “Automated Testing of COBOL to Java Transformation.” Type: IBM industrial-experience preprint · Published: 14 April 2025 · Version: arXiv:2504.10548v1 · Pinpoint: Abstract and introduction · Accessed: 15 August 2026 · Supports: manual validation remains necessary in the described transformation workflow. Source

[32] Kahn, Probasco & Kinoshita, AI Safety and Automation Bias. Type: policy-research report with three case studies · Published: November 2024 · Version: CSET report, DOI 10.51593/20230057 · Pinpoint: Executive Summary and conclusion · Accessed: 15 August 2026 · Supports: human-in-the-loop is not sufficient by itself. Source

[33] Perry et al., “Do Users Write More Insecure Code with AI Assistants?” Type: peer-reviewed user study · Published: ACM CCS 2023; full version revised 18 December 2023 · Version: arXiv:2211.03622v3 / CCS '23 pp.2785–2799 · Pinpoint: Abstract; §§4–5 · Accessed: 15 August 2026 · Supports: security outcomes and participant confidence in the studied setting. Source

[34] PREPRINT: NOT PEER REVIEWED. Rosbach et al., “When Two Wrongs Don't Make a Right.” Type: controlled-study preprint · Published: 1 November 2024 · Version: arXiv:2411.01007v1 · Pinpoint: Abstract; experiment with 28 trained pathology experts · Accessed: 15 August 2026 · Supports: false AI confirmation can reinforce an already-wrong judgment in that domain. Source

[35] PREPRINT: NOT PEER REVIEWED. Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.” Type: randomized controlled trial preprint · Published: submitted 12 July 2025; revised 25 July 2025 · Version: arXiv:2507.09089v2 · Pinpoint: Abstract; Table 2; core result · Accessed: 15 August 2026 · Supports: task population, forecasts, and 19% measured slowdown in that setting. Source

[36] PREPRINT: NOT PEER REVIEWED. Peng et al., “The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.” Type: controlled-experiment preprint · Published: 13 February 2023 · Version: arXiv:2302.06590v1 · Pinpoint: Abstract; experimental design and result · Accessed: 15 August 2026 · Supports: 55.8% faster completion on the specified JavaScript task. Source

[37] Google Cloud DORA, Accelerate State of DevOps Report 2024. Type: industry survey and statistical modelling report · Published: October 2024 · Version: 2024 report · Pinpoint: “Impact of Generative AI in Software Development” — an estimated 1.5% reduction in throughput and 7.2% reduction in delivery stability for every 25% increase in AI adoption, visualised in Figure 10 · Accessed: 21 August 2026 · Supports: associations with throughput and delivery stability; not causal proof. Source

[38] Flyvbjerg & Budzier, “Why Your IT Project May Be Riskier Than You Think.” Type: management-research article using a 1,471-project dataset · Published: September 2011 · Version: Harvard Business Review 89(9); author manuscript archived 2013 under the variant title “Why Your IT Project Might Be Riskier Than You Think” · Pinpoint: sample results on mean and tail overruns · Accessed: 21 August 2026 · Supports: material tail risk hidden by point estimates. Manuscript

[39] NIST/SEMATECH, e-Handbook of Statistical Methods, §§7.2.4.2 and 6.2.6. Type: official statistical handbook · Published: web edition, undated revisions · Version: live web edition at access date · Pinpoint: one-sided sample size for a proportion; sequential-sampling acceptance and rejection limits · Accessed: 18 August 2026 · Supports: fixed-sample planning formula, continuity correction, and α/β sequential boundaries. Sample size · Sequential plan

[40] Eldridge, Ashby & Kerry, “Sample size for cluster randomized trials: effect of coefficient of variation of cluster size and analysis method.” Type: peer-reviewed methodological study · Published: 30 August 2006 · Version: International Journal of Epidemiology 35(5), 1292–1300; DOI 10.1093/ije/dyl129 · Pinpoint: unequal-cluster design-effect approximation and analysis qualifications · Accessed: 18 August 2026 · Supports: inflation for within-cluster correlation and unequal cluster sizes. DOI

[41] Goel, Struber, Auzina, Chandra, Kumaraguru, Kiela, Prabhu, Bethge & Geiping, “Great Models Think Alike and this Undermines AI Oversight.” Type: peer-reviewed empirical study; ICML 2025 spotlight poster · Published: 6 February 2025 · Version: ICML 2025; arXiv:2502.04313 · Pinpoint: Abstract; CAPA definition and the judge-similarity and capability-correlation results · Accessed: 20 August 2026 · Supports: a model judge favours models similar to itself, and model mistakes grow more similar as capability increases. Source

[42] PREPRINT: NOT PEER REVIEWED. Verga, Hofstätter, Althammer, Su, Piktus, Arkhangorodsky, Xu, White & Lewis, “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.” Type: evaluation-method preprint; indexed by dblp as a CoRR record with no proceedings entry located · Published: 29 April 2024 · Version: arXiv:2404.18796v2 · Pinpoint: Abstract; panel composition of disjoint model families and the reduction in intra-model bias · Accessed: 20 August 2026 · Supports: a judge panel reduces single-judge bias when its members come from disjoint model families. Source

[43] Knight & Leveson, “An Experimental Evaluation of the Assumption of Independence in Multiversion Programming.” Type: peer-reviewed journal article · Published: January 1986 · Version: IEEE Transactions on Software Engineering 12(1), pp.96–109; DOI 10.1109/TSE.1986.6312924; author-archived full text read · Pinpoint: Abstract; §5 model of independence and the 99%-confidence rejection of the independence assumption; conclusions · Accessed: 20 August 2026 · Supports: twenty-seven versions written independently from one specification, subjected to one million tests, failing coincidentally substantially more often than independent failure predicts. Author copy · DOI

[44] PREPRINT: NOT PEER REVIEWED. Brown et al., “Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.” Type: research preprint · Published: submitted 31 July 2024; revised 30 December 2024 · Version: arXiv:2407.21787v3 · Pinpoint: Abstract; coverage scaling over four orders of magnitude and the no-verifier plateau of majority voting and reward models · Accessed: 20 August 2026 · Supports: repeated sampling converts into delivered performance where answers can be automatically verified; selection without an automatic verifier plateaus beyond several hundred samples. Source

[45] Ding et al., “Vulnerability Detection with Code Language Models: How Far Are We?” Type: peer-reviewed empirical study · Published: ICSE 2025; preprint revised 10 July 2024 · Version: ICSE 2025 proceedings pp.1729–1741, DOI 10.1109/ICSE55347.2025.00038; arXiv:2403.18624v2 · Pinpoint: Abstract; PrimeVul design (label quality, de-duplication, chronological split, paired vulnerable/patched functions, §IV-B2) and the BigVul-to-PrimeVul F1 results · Accessed: 20 August 2026 · Supports: 68.26% to 3.09% F1 for a state-of-the-art model under corrected evaluation design, with GPT-3.5/GPT-4 attempts akin to random guessing in the most stringent settings. Source · DOI

[46] PREPRINT: NOT PEER REVIEWED. Miller, “Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.” Type: methodology preprint · Published: 1 November 2024 · Version: arXiv:2411.00640v1 · Pinpoint: §2 analysis framework (variance decomposition into question-selection and conditional components); §3.1 resampling · Accessed: 20 August 2026 · Supports: resampling one input reduces only the expected conditional-variance component of evaluation variance; it adds no evidence across inputs. Source

[47] Nancy G. Leveson, Engineering a Safer World: Systems Thinking Applied to Safety. Type: open-access scholarly monograph · Published: 2011 · Version: MIT Press open-access edition, ISBN 978-0-262-29824-7 · Pinpoint: chapter 2, pp.7–15; chapter 3, pp.61–100; chapters 7–8, pp.181–220; chapters 10–13, pp.307–444; chapter 14, pp.445–461; Appendix B, pp.469–493 (illustrative case only) · Accessed: 25 August 2026 · Supports: system properties arising from interactions; constraint, controller, action, feedback, and process-model analysis; four unsafe-control-action categories; assumption tracking; coordination analysis; and lifecycle learning. Transfer boundary: the modernization schema in §8.1 is this paper's adaptation; the source does not establish modernization efficacy, regulatory compliance, or a complete modernization risk model. Open-access PDF

Appendix B — Selected sample evidence artifacts

Status date: 16 August 2026. SAMPLE DATA—NOT CLIENT EVIDENCE. This public appendix contains the artifact catalogue, one reduced coverage example, and the responsibility matrix. It shows how to name a gap and assign authority. It does not publish implementation schemas, provenance contracts, thresholds, or populated operating registers. Blind calibration, full comparison, reconciliation, gate-ledger, and sensitivity-record outputs remain required even though they are not reproduced here.

Artifact Required source Boundary of the sample
Coverage map Client instrumentation plus historical and synthetic evidence The method never infers behavioural coverage from extract size; every uncovered surface remains named
Divergence budget Human-approved tolerances declared before comparison Sample tolerances are illustrative and cannot be adopted as client thresholds
Intent register Human decisions linked to divergences and approval evidence The record preserves accountability; it cannot decide whether a behaviour should change
Monthly drift report Authoritative release, parameter, corpus, and calibration records Automation may join and age records; an authorised reviewer owns semantic impact
One-way-door register Human-authored inventory of effects, compensation, owners, and authorisers No tool can discover the business effect, invent compensation, or authorise release
RACI Human-assigned roles resolved to controlled identities A responsibility matrix states authority; separate operating evidence must show that the control acted

B1. Reduced coverage example

This deliberately non-normative example shows the public decision shape without publishing the production schema or provenance envelope.

Behavioural surface Observed sample population Coverage finding Required action
Daily interest accrual July accounts, one historical leap-day replay, and synthetic rate-boundary cases Normal and negative cases observed; rate boundaries remain synthetic-only Obtain live boundary evidence before scale
End-of-day posting 31 production cycles Happy path and retry observed; late reversal absent Capture a reversal cycle; scale gate remains blocked
Annual tax statement No event in the observation window UNEVIDENCED Run a synthetic cycle and capture the next live annual event

Required rule: an uncovered path is named. It is never converted into a zero-divergence claim. Synthetic cases remain labelled and state-bound, are reported separately from observed production coverage, and cannot establish production incidence or satisfy a live-evidence requirement.

B2. RACI

R = Responsible · A = Accountable (exactly one per row) · C = Consulted · I = Informed

Activity Sponsor Business decision owner Legacy SME Modern engineering Independent reviewer Data owner Risk / compliance Procurement / legal Cut-over authority
Approve programme shape and post-cut-over value plan A R C C I C C C I
Approve collection boundary, access, retention, and deletion I C I R C A C C I
Confirm jurisdiction, covered entity/operator, and system classification I C C I C C A/R C I
Capture and classify legacy behaviour I A R R C C C I I
Approve divergence, service, security, change, and freshness budgets I A C R C C C C I
Adjudicate deliberate behaviour change I A R C C C C I I
Validate evidence and independence I C C I A/R C C I I
Approve data reconciliation I C C R C A C I I
Approve commercial off-ramps and maximum committed spend A C I C C C C R I
Authorise first irreversible external write I C I R C C C I A
Accept pilot stop, reshape, or scale decision A R C C C C C C C

The migration author cannot fill the independent-review role for the same work. RACI names roles. The controlled assignment register links each gate role to a person. A missing assignment renders the gate UNEVIDENCED. The irreversible-write accountable role is a named human, not an agent or generic team.

Appendix C — Glossary

Term Meaning in this paper
AI/ML Artificial intelligence / machine learning. Regulatory texts may define these terms more narrowly than ordinary usage.
ARITH(EXTEND) Enterprise COBOL compiler option enabling extended precision for supported arithmetic; relevant to decimal equivalence.
AST Abstract syntax tree: a structural representation of source code. AST diffs can nominate changed logic but do not prove complete semantic impact analysis.
Behavioural/oracle coverage Share of declared observable surfaces and populations for which versioned legacy evidence exists; not source-line coverage.
CDC Change Data Capture. A family of techniques for observing data changes; the safe mechanism and its coverage depend on the platform and workload.
Change budget Preapproved count or scope of deliberate behaviour changes allowed during the conversion, separate from the defect/noise budget.
Change-port protocol Controlled process for carrying a legacy production change into a different target language or architecture through linked intent, impact, implementation, and fresh evidence—not textual cherry-picking.
CICS IBM Customer Information Control System, a mainframe transaction-processing environment.
COBOL Common Business-Oriented Language, widely used in long-lived business systems.
COMMAREA CICS communication area used to carry application state between pseudo-conversational tasks.
COMP-3 COBOL packed-decimal representation: two decimal digits per byte plus a sign nibble.
Divergence budget Preapproved tolerance by divergence class, fixed before comparison results are known.
EBCDIC IBM mainframe character encoding whose collation differs from ASCII/UTF-8.
FRFI Federally Regulated Financial Institution in Canada; the population to which OSFI guidance may apply according to scope.
Freshness budget Maximum permitted age or lag between legacy change and recaptured/recalibrated evidence.
Golden master Versioned recorded legacy behaviour used as expected evidence for captured conditions; not proof that the behaviour is correct or desirable.
Human separation Separation between the people who develop migration work and those who independently review or authorise it.
Intent register Decision record for each deliberate behaviour change: what changed, why, who approved it, when, and against which evidence/budget.
IRB Internal Ratings-Based approach used by authorised banks to calculate regulatory credit-risk capital; several model-validation sources in §2 are IRB-scoped.
LPAR Logical partition: an isolated operating environment on a mainframe.
Machine separation Independent provenance for judging evidence and thresholds, followed by protection from later change by the migration author; a control inference in this paper, not quoted regulation.
MIPS Millions of Instructions Per Second; a capacity/performance proxy, not the traditional sub-capacity billing determinant described in §1.3.
MLC Monthly License Charge for eligible IBM Z software.
MRM Model Risk Management.
One-way door An effect that code rollback cannot undo, especially an external or non-idempotent write.
Operational envelope The declared workload conditions and budgets for latency, throughput, concurrency, resource consumption, saturation, and recovery.
Oracle A source that determines expected behaviour. Here, recorded legacy execution is a practical oracle, not proof that legacy behaviour is desirable.
Parameter drift Change to rates, tables, thresholds, calendars, or other runtime data that alters behaviour without a code change.
pass@1 Proportion of tasks solved by the first generated candidate under the benchmark's stated success test.
PBI Product Backlog Item: a unit of planned product or engineering work.
Pseudo-oracle An independently produced alternative implementation used for comparison when a complete expected-output oracle is unavailable.
R4HA Rolling four-hour average. In the cited IBM pricing model, the monthly peak R4HA drives traditional sub-capacity charges.
RACI Responsibility matrix: Responsible, Accountable, Consulted, Informed. This paper requires exactly one accountable role per activity.
RBA Relative byte address used to access a VSAM record by physical position; not itself an access method.
Replay corpus Versioned set of recorded inputs, outputs, parameters, and boundary effects used for comparison.
Rule drift Change to executable business logic that can make previously captured evidence stale.
Shadow execution Running the candidate on the same input as production while isolating its side effects, then comparing outputs.
SPRT Sequential Probability Ratio Test. An optional sequential hypothesis test whose rates, error bounds, sampling assumptions, and stopping rule must be predeclared; not a universal equivalence gate.
VSAM IBM Virtual Storage Access Method for mainframe datasets.
xfail Expected-failure marker in test frameworks. New or unexplained markers are evidence-integrity findings because they can hide failures.

Appendix D — Evidence gaps and excluded claims

D1. Named evidence gaps

  1. No located empirical work measures LLM translation of Delphi or PowerBuilder at repository scale.
  2. No published repository-scale COBOL migration success rate was located. COBOL evidence is smaller-grained. Repository evidence is primarily Java.
  3. No peer-reviewed post-deployment study tracks production incident rates in systems migrated from legacy languages by LLM.
  4. No real-world causal taxonomy of modernization failures provides defensible base rates by cause.
  5. No large-N outcome comparison of big-bang versus incremental migration was located. The preference for incremental migration rests mainly on risk structure and practitioner evidence.
  6. No study located quantifies dual-maintenance cost as a portable FTE or dollar benchmark. Use client change history.
  7. No cross-organisation dataset supports a claim that modernization programmes “typically” take a particular number of years.
  8. Golden-master decay has not been studied specifically for migrations. §6 infers the mechanism from differential-testing and test-co-evolution evidence.
  9. Feature-freeze evidence is interview-based, not a measured rate of attempted or sustained freezes.
  10. No located study measures whether separating AI authorship of code from AI authorship of tests reduces escaped defects. §2.4 argues the mechanism from evaluation research and reports no effect size for this control.
  11. No located outcome study tests the §8.1 modernization adaptation as a complete intervention. Its fields are a proposed control design, not an observed effect or portable assurance rate.

D2. Considered and not used

  • “68–79% of modernization projects fail or underperform”: no traceable primary source located.
  • Standish CHAOS success/failure rates: categories and methods are unsuitable for the claimed use.
  • Gartner forecasts about mainframe-exit failure: forecasts without a published sample are not measurements.
  • Vendor case-study fidelity percentages: excluded where population, denominator, and comparison method are undisclosed.
  • Dual-running cost and outage-cost figures: excluded where repeated through consultancy pages without an attributable measurement.
  • “Modernize-in-place succeeds 71% versus 26% for rewrite”: vendor-commissioned and inaccessible primary methodology. Excluded despite supporting the paper's position.
  • Specific driver percentages surfaced by search summaries but absent from the cited practitioner report: excluded.

D3. Citation and inference controls

  • The EU Capital Requirements Regulation is not cited for development-versus-validation separation. The paper uses the ECB guide within its stated scope.
  • FINMA 08/2024 §2.7 is an observed supervisory finding, not a legislative requirement.
  • Singapore's 2024 information paper is good-practices material, not binding AI-specific rules, and is not used as a mandate here.
  • The machine-separation argument is this paper's synthesis. No cited regulator states that AI code authorship and AI test authorship must be separated. The supporting evidence in §2.4 comes from language-model evaluation research and the multi-version programming experiment. It does not come from migration research. The paper cites it for the mechanism, not for a transferable rate. [41] [42] [43]
  • No located study measures correlated model error in legacy-code migration specifically. The provenance rule in §2.5 is therefore a control design proposed here, argued from evaluation evidence and from the structure of the failure in §4.
  • The cross-language change-port protocol and prohibition on semantic auto-merge are control designs proposed in this paper.
  • The composition tables in §8 extend an illustration the IBM paper states only for a ten-paragraph COBOL program at 90% accuracy per paragraph. Every other row is computed here. So is the inverted per-unit accuracy table. These computations use an independence assumption that the source does not state and this paper does not endorse.
  • Fowler supports incremental value and the ability to retain value if a strangler programme stops partway. The rule that transitional architecture belongs in the estimate is this paper's inference, not source wording.
  • Section 8.1 transfers systems-theoretic control concepts from Leveson into modernization. The outcome-to-state-to-constraint chain, shared control record, four action-failure prompts, seam review, deployment binding, and maintained-approval cycle are proposed control designs. They are not regulator-authored requirements, a safety case, or evidence that the method works. [47]
  • Sizing the sensitivity pack by the zero-event binomial bound is standard statistics applied here to gate sensitivity. The application is this paper's, but the arithmetic is not. The distinct-inputs counting rule is part of the application: the bound's independence assumption is between seeded inputs, not between executions of one input. Semantic diffs can nominate impact. They do not prove completeness or equivalence.
  • The commercial off-ramp provisions in §12.1.1 are minimum control-design recommendations, not jurisdiction-specific legal advice. Client procurement and legal functions must adapt them.
  • The iTnews citation uses the exact article permalink, not the publication homepage. The article directly quotes CBA's CEO at the 2012 AGM. [23]

The delivery method built on this argument is written up as The PROVEN Migration Methodology.

← Back to white papers  ·  Read the executive brief