Executive brief
The Oracle You Already Own — executive brief
Pierre Boutquin · Version 1.8 · 3 September 2026 · first published 25 August 2026 · Audience: CIOs, CTOs, engineering leaders, business owners, risk, audit, and model-governance stakeholders
Three questions, in order
- Is it complete? Declare the scope and denominator first. Account for every in-scope behaviour, rule, record, state, parameter, and external effect as evidenced, missing, excluded with authority, or UNEVIDENCED. A fully accounted assessment may still answer no. Required material evidence must be present before this gate closes.
- Is it consistent? Once the completeness gate closes, compare old and new behaviour, service envelope, data, and security posture. Remediate every inconsistency or record an authorised deliberate departure. Any unexplained critical break or unapproved budget breach blocks the sequence.
- Where does it say that? Once each item has a canonical disposition, publish the source, comparison, decision, owner, date, provenance, and freshness record. That record is the deed. It records the authorised result rather than creating truth by assertion.
The programme instruments controls from the start. It answers the questions in sequence. It captures provenance and freshness before closeout because they cannot be reconstructed honestly.
Regulated overlay
For regulated entities, two outer gates wrap the universal sequence:
- What applies? Establish the jurisdiction, instrument, legal force, entity, role, system, obligation population, and effective date. An adjacent control does not become binding because it sounds sensible.
- Is it compliant? A qualified human decision-maker acts after the evidence is complete, consistent, and traceable. The decision-maker assesses every obligation in the applicability register. The four dispositions are satisfied, not satisfied, UNEVIDENCED, and not applicable—with authority and rationale.
A traceable test result is not a compliance determination.
The permitted positive verdict is compliant for the declared regulatory scope, as of the stated date. A generated report supports that judgment. It is not a standalone conformity assessment, legal opinion, or cut-over authorisation.
The decision
Approval requested: authorise a bounded evidence pilot covering governance, capture, calibration, budget design, and non-mutating replay. Do not authorise estate-wide conversion or cut-over. The charter fixes scope, duration, spend, accountable executive, risk appetite, and excluded irreversible effects. Build, scale, and live-write authority remain later decisions.
The sponsor must also fund the possibility that the answer is “no”. That funding includes a protected reviewer, a pre-agreed escalation path, decision time, and schedule and cost reserve. It applies when one unresolved semantic difference blocks a cut-over. A reviewer punished for using the authority in the governance chart is not a gate.
A like-for-like conversion buys no new capability. It buys the option to change a system that is expensive, slow, risky, or difficult to integrate. Before approval, identify the post-cut-over changes that will exercise that option, with an owner and date for each. Without that plan, compare re-implementation, package replacement, encapsulation, and retention. Age alone is not a business case. The charter must quantify benefits, dual-run and exit costs, timing, downside, and who will realise the benefits.
For IBM Z, IBM's mainframe platform, some estates use traditional sub-capacity Monthly License Charge (MLC) pricing. The charges track the highest observed four-hour rolling average (R4HA). Therefore, infrastructure savings accrue only if removal lowers the monthly peak R4HA. Lower average CPU use is not the billing test. [5]
The sponsor must explain what the conversion makes possible and who will use it. The delivery team must show how it will detect a different answer from the new system. When a difference appears, a named business authority must decide which behaviour should govern future operation. Each answer needs an operating mechanism. The investment case and evidence gates also need numbers.
Before capture, require signed scope and funding. Require client-specific coverage, calibration, divergence, freshness, and change budgets. Require data-owner approval for collection, access, retention, deletion, and pseudonymisation. Require risk and compliance sign-off on jurisdiction and system classification. Also require a costed capability plan that separates shipped tools from manual controls. Meter rule re-porting, parameter re-evaluation, and carrying load from target-release freezes separately.
The risk
Translation can be fast while proof remains difficult. In ScarfBench, five coding agents attempted 204 cross-framework Java migrations. The strongest result was 15.3% aggregate test pass on focused migrations and 12.2% on whole applications. One task was fully equivalent. This is cross-framework evidence and should not be treated as a COBOL-to-Java proxy.† [1]
AgentModernize reports a 23.0% mean behavioural equivalence rate on eight synthetic scenarios. That value is a success rate and evidence of difficulty in the study, not an error rate or an industry baseline.† [2]
The central control is independent behavioural evidence captured before the old system disappears. The legacy system is not automatically right, but it is the available witness to what production did. Use fidelity as the default treatment for behaviour that nobody has reconsidered. Every deliberate difference needs a business owner, rationale, evidence, date, and approved change budget.
The evidence pack
Do not accept a single “equivalence percentage” as sufficient evidence. Require six measurements grouped under the three ordered questions, each with denominators, eligible counts, and explicit gaps:
| Order |
Question |
Measurements |
Exit condition |
| 1 |
Is it complete? |
1. Oracle coverage · 2. Comprehension calibration |
Scope and denominator are explicit. Every item is accounted for. The assessment can answer no, and missing required material evidence keeps the gate open. |
| 2 |
Is it consistent? |
3. Behaviour and service divergence · 4. Data and security posture |
Functional, operational, data, and security results pass separately. No critical break or unapproved budget breach remains. |
| 3 |
Where does it say that? |
5. Gate integrity · 6. Drift and freshness |
Every result resolves to protected, current evidence and a person-linked decision record. |
For regulated entities, the applicability register defines the obligation population before these six measurements begin. After they close, an obligation matrix records the compliance disposition and evidence locus for every applicable obligation. The six measurements establish the evidence record. They do not turn it into an unqualified compliance verdict.
| Measurement |
Executive question |
Required artifact |
| 1. Oracle coverage |
Which behaviours and populations were actually observed? |
Coverage map |
| 2. Comprehension calibration |
Can the method recover rules whose answers are already known? |
Blind calibration report |
| 3. Behaviour and service divergence |
How often do old and new differ by class, and does the target meet predeclared latency, throughput, concurrency, resource, saturation, and recovery budgets? |
Behavioural and service budgets + comparison report |
| 4. Data and security posture |
Do records, aggregates, encodings, and accounting identities reconcile, and does the target's threat model and exposed surface satisfy approved security requirements? |
Reconciliation report + security assessment |
| 5. Gate integrity |
Is judging evidence independently sourced and protected from later change? Is human review separate, and can the gate reject every applicable seeded-defect class? |
Gate ledger + responsible, accountable, consulted, and informed (RACI) responsibility matrix + sensitivity record |
| 6. Drift and freshness |
Is the evidence keeping pace with changes to rules and parameters? |
Monthly drift report |
Add two decision controls. Use an intent register for deliberate behaviour changes. Use a one-way-door register for writes that code rollback cannot undo. Functional agreement cannot absorb a service-envelope breach. Data reconciliation cannot absorb a security finding. Each result keeps its own population, budget, owner, and stop condition.
A passed component is not yet a controlled system. For every consequential gate, record:
- the unacceptable outcome, prohibited system state, and enforceable constraint.
- the accountable controller, stopping authority, required actions, and forbidden actions.
- current assumptions, evidence feedback, maximum useful delay, and hand-off gaps or conflicts.
- the exact deployed subject and the events that force revalidation.
Challenge four failure modes:
- omission
- an unsafe action
- the wrong timing or sequence
- an action that ends too soon or persists too long
This is a modernization control design, not a safety, compliance, or efficacy claim.
Any sampled rate declares acceptable and intolerable rates, and , the independent sampling unit, batch clustering, and the stopping rule before results are seen. Volume alone does not establish independence. Synthetic cases remain separately labelled and state-bound. They may challenge uncovered paths but cannot estimate production incidence or supply their own expected answers.
The public examples in the full paper show how to name a coverage gap and assign responsibility. They are not client evidence or implementation schemas. Before the programme uses a gate, the pilot charter identifies automated and manual controls, their owners, independent QA, and cost. Planned or illustrative capability is not present evidence.
The regulatory boundary
Do not state that every migration is a model modification. OSFI E-23 will apply only when the system meets its model definition. In scope, relevant changes trigger review independent of development. It was published on 11 September 2025 and takes effect on 1 May 2027. At the paper's status date, it is a preparation benchmark rather than an effective requirement. [3]
Other supervisory instruments reinforce independent challenge only in their own scopes. They do not create a universal modernization rule. Machine separation between AI-produced code and acceptance evidence is this paper's inference, not regulator-authored text.
For providers of covered high-risk AI systems, the EU AI Act requires applicable third-party development tools in technical documentation. Relevant duties apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I product systems. It does not require an AI-authored percentage, and a coding assistant is not automatically high risk. [4]
Accordingly, “What applies?” is an explicit gate, not a footnote. The programme records the entity, role, system, jurisdiction, instruments, obligations, legal force, and effective date. Only then can it ask whether complete, consistent, traceable evidence satisfies each applicable duty.
The pilot decision
Design the pilot to expose failure cheaply. Publish and contract the stop, scale, and off-ramp criteria before work starts. A breached gate stops the current path and does not authorise a looser threshold.
- Reshape a bounded, non-critical evidence gap under a versioned new charter, prospective budgets, and a fresh holdout or replay window.
- Pivot when direct conversion does not converge or platform coupling makes it the wrong architecture. Evaluate encapsulation, re-implementation, or package sourcing under new gates.
- Stop / retain when critical reconciliation, evidence integrity, authority, or the investment case fails. End conversion, apply the approved evidence-retention plan, and assign the safe-operating or replacement decision.
Contract for a client-runnable evidence path. Give the client custody and usable rights for:
- the corpora, adapters, schemas, and normalisation rules
- the thresholds, fixtures, and reports
For the harnesses, give the client either:
- the replay, divergence, service, and security harnesses
- the source, build instructions, dependency pins, runbooks, and interfaces needed to reproduce their outputs in its own controlled environment
A provider-hosted black box cannot be the sole path to the evidence supporting the decision.
Proceed only when the programme reports all six measurements within client-approved budgets. Show denominators, eligible counts, and gaps. Each relied-on gate must have rejected every required seeded-defect class that applies. The record must show each rejection. Classify and accept remaining divergences by name. Service and security results must pass as separate strata. The intent register needs accountable owners, and data and regulatory approvals must be current. Bind the evidence to the executable, configuration, parameters, data contract, environment, and release that will operate. Assign revalidation triggers and effectiveness checks. Non-idempotent or external effects need person-linked human authorisation. Each claimed reversible effect needs tested compensation. The release needs a rehearsed fix-forward package for rollback or compensation failure. Where regulation applies, also require a current applicability register. Before a positive compliance claim, the obligation-by-obligation assessment must mark no applicable duty as not satisfied or UNEVIDENCED.
Version control makes code rollback routine. However, some effects can outlive the code that produced them:
- a new identifier
- a posted transaction
- a settlement instruction
- a regulatory filing
- a customer notification
Therefore, the first irreversible external write requires release-specific human authorisation. The release also needs a rehearsed fix-forward path. It covers traffic isolation, known-good deployment, state or schema repair, reconciliation, evidence refresh, decision rights, and recovery time. This path is necessary because rollback or compensation can fail during a live incident.
What approval buys
The pilot produces:
- A captured oracle.
- A calibrated method.
- Published behavioural and service budgets.
- Contracted off-ramps.
- A client-runnable harness.
- A measured comparison.
- Data reconciliation.
- A security assessment.
- Intent and drift records.
- A one-way-door register.
- A rehearsed fix-forward package.
- Signed accountability.
This evidence supports:
- A funded reshape.
- Encapsulation.
- Replacement.
- Retention.
- A decision not to modernize this system by this method at its current readiness.
Numbered source notes
† Preprint, not peer reviewed. Submitted but not yet through peer review. It is cited as the best available evidence on the point and marked because it carries less evidential weight than a peer-reviewed or official source.
[1] PREPRINT, NOT PEER REVIEWED. Pavuluri et al., “ScarfBench.” Type: research preprint · Published: submitted 7 May 2026; revised 18 May 2026 · Version: arXiv:2605.06754v2 · Pinpoint: Abstract; §5.1; Tables 3–4 · Accessed: 15 August 2026. https://arxiv.org/abs/2605.06754
[2] PREPRINT, NOT PEER REVIEWED. Ahmed & Galib, “AgentModernize.” Type: research preprint · Published: submitted 17 May 2026; revised 4 August 2026 · Version: arXiv:2605.17535v2 · Pinpoint: Abstract; Tables 3–4; §VI · Accessed: 15 August 2026. https://arxiv.org/abs/2605.17535
[3] OSFI, Guideline E-23: Model Risk Management (2027). Type: official supervisory guideline · Published: 11 September 2025; effective 1 May 2027 · Version: final E-23 · Pinpoint: A.4 “Model”; Principle 3.4 “Model review” · Accessed: 15 August 2026. https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/guideline-e-23-model-risk-management-2027
[4] Regulation (EU) 2024/1689, Artificial Intelligence Act, as amended by Regulation (EU) 2026/1744. Type: primary legislation · Published: original 12 July 2024; amendment 24 July 2026 · Version: Official Journal texts in force at 15 August 2026 · Pinpoint: Articles 6, 17 and 113; Annexes I, III and IV §2(a); amending Regulation Article 1(40) · Accessed: 15 August 2026. AI Act · 2026 amendment
[5] IBM, z/OS 3.1: Planning for Sub-Capacity Pricing. Type: official vendor manual · Published: updated 17 May 2024 · Version: SA23-2301-60 · Pinpoint: “Traditional sub-capacity pricing,” definition of highest observed four-hour rolling average · Accessed: 15 August 2026. PDF
Read the full paper →
The delivery method built on this argument is written up as The PROVEN Migration Methodology.