Why Banks Still Run Mainframes, and What Replacing One Actually Costs

Why Banks Still Run Mainframes, and What Replacing One Actually Costs

Banks still run mainframes because, for the job the mainframe does, high-volume transaction processing that must never lose or duplicate a record, it remains very good. The problem is not the machine. It is the decades of undocumented business rules running on it, the shrinking pool of people who can read them, and a pricing model that makes growth expensive. Executives are usually offered two stories: the mainframe is a relic to be retired, or it is irreplaceable and should be left alone. Both are wrong in ways that cost money. Mainframe modernization programs fail when they treat the problem as a code conversion, and succeed when they treat it as a risk-managed migration of data, rules and operations, one slice at a time.

Four paths

Key Takeaways

  • The mainframe runs the ledger and the nightly batch cycle, the two parts of a bank where a lost or duplicated record is unacceptable. Its throughput and reliability are genuine, which is why replacement is hard to justify on technical grounds alone.
  • The code is not the hard part. The hard part is the business rules embedded in it, many written decades ago, some contradicting each other, and few documented anywhere except in the code itself.
  • The cost of staying has three parts: capacity-based licence pricing that rises with peak usage, a skills pool that is retiring faster than it is replaced, and the risk cost of rules nobody can confidently change.
  • Big-bang replacements carry the highest risk and the most public failures. Programs that succeed migrate incrementally, run old and new in parallel, and move data before moving logic.
  • AI code tools are genuinely useful for understanding and documenting a COBOL estate. Automatic rewrite into a modern language without extensive testing is where they are oversold.
Big bang slices

What a Mainframe Actually Does in a Bank

A mainframe is a large, highly redundant server designed to process very large volumes of transactions with near-continuous availability and strict data integrity. In a bank it typically runs two kinds of work.

Online transaction processing. Every card authorization, transfer and balance inquiry that touches the core ledger arrives as a transaction that must be processed once, completely, and in order. Mainframe transaction monitors and databases were built for exactly this, at volumes of thousands of transactions per second with integrity guarantees that are part of the platform rather than something each application must engineer.

Batch processing. Overnight, the bank runs the jobs that turn a day of transactions into a closed set of books: interest accrual, fee calculation, statement generation, regulatory reporting, and the files exchanged with payment networks and clearing houses. This nightly cycle is sequenced and interdependent, and if one job fails, later jobs wait. Much of what customers experience as "pending until tomorrow" is the batch window.

The ledger is the center of gravity. As covered in core banking systems, it is the authoritative record of every account, and everything else in the bank, from mobile apps to fraud systems, reads from it or writes to it.

The scale of the dependency is large, though the widely quoted figures are dated and partly vendor-sourced. IBM states that its Z mainframes handle over 70 percent of global transactions by value, a claim drawn from research IBM commissioned. A 2017 Reuters analysis estimated that 220 billion lines of COBOL were still in use and that COBOL underpinned 43 percent of banking systems. Both numbers are repeated widely and neither has been independently refreshed, so they are best read as indicating direction rather than current magnitude.

Ai code tools

Why It Has Survived

Three reasons, in increasing order of importance.

Throughput and reliability. For many small, integrity-critical transactions, the mainframe is fast, reliable and well understood. Distributed systems can match it, but only by engineering for what the mainframe provides by default.

The cost of proving equivalence. Replacing a system that works requires proving the new one produces the same results for every product, every edge case and every historical account state. That proof is expensive, and the benefit of replacement is often indirect.

Business rules nobody has written down. This is the decisive reason. A core banking estate accumulates rules over decades: how interest is calculated on a product discontinued in the 1990s, how a particular rounding convention was agreed with a regulator, what happens when a transfer crosses midnight on the last day of a leap year. Many rules exist only in code. Some were written by people who have retired. Some contradict each other and are resolved by the order in which jobs run. The mainframe is the only complete specification of how the bank behaves, and that is why replacing it is a knowledge-recovery problem before it is an engineering one.

The Real Cost Structure

Executives usually see the mainframe as a licence and hardware line. The full cost has three components, and the third is the one that decides strategy.

Capacity pricing

Mainframe software is commonly priced on capacity rather than on users. Under IBM's monthly licence charge model, charges for many products are driven by the peak rolling four-hour average of processing capacity consumed in a month. The effect is that a bank pays for its busiest hours, and any new workload that raises the peak, such as a mobile app that multiplies balance inquiries, raises the software bill. IBM and other vendors now offer alternative consumption-based pricing models, and negotiating these is often the fastest cost saving available, but the underlying dynamic, paying for peak, is what makes growth on the mainframe expensive.

A common optimization is to move read-heavy work off the mainframe entirely, replicating ledger data to a modern data store that serves balance inquiries and analytics, so the mainframe processes only the transactions that change the ledger.

The skills pool

The engineers who built and understand these estates are concentrated in older age groups, and fewer new engineers learn COBOL, mainframe job control or the transaction monitors. The risk is not that nobody can write COBOL, which is learnable, but that the people who know why a program behaves as it does are leaving.

The risk cost of unreadable rules

This is the cost that rarely appears on a budget line and most often drives the decision. When nobody can confidently predict what a change will do, every change is slow and expensive: extended testing, change freezes, and a growing list of products the bank cannot modify. That is technical debt in its most literal form, and it can be measured by change lead time and failure rate, as set out in measuring technical debt. A bank that cannot launch a new product for eighteen months because the core cannot safely be changed is paying a cost far larger than its licence bill.

The Four Modernization Paths

Rehost, or lift and shift

Move the existing applications, largely unchanged, from mainframe hardware to an emulated or re-platformed environment on commodity servers or cloud infrastructure. The COBOL stays COBOL.

This reduces hardware and some licence cost and can be done relatively quickly. It does not reduce the knowledge problem at all, since the same undocumented rules now run somewhere else. It suits institutions whose main pain is cost or data center exit rather than agility.

Refactor and code conversion

Translate the COBOL into a modern language, automatically or semi-automatically, preserving behavior. This is where AI-assisted translation tools are now marketed heavily.

Conversion produces code in a modern language, but converted code often preserves the structure of the original, producing a modern-language program that is as hard to understand as the COBOL it replaced. The testing burden is the real cost: every converted program must be shown to behave identically, including on edge cases nobody remembers. It suits well-bounded applications with good test coverage, or where the conversion is followed by genuine redesign.

Strangler pattern around the core

Build new capabilities outside the mainframe and gradually route functions away from it, one product or process at a time, until the remaining core is small enough to replace or retire. The name comes from Martin Fowler's description of a strangler fig that grows around a host tree. Typically, the first step is an API layer in front of the mainframe, so new channels talk to the API rather than directly to the core.

This spreads risk across many small, reversible migrations and delivers value early, at the price of running two estates for years. The architectural choices involved, including when to break functions into services and when a modular design is enough, are discussed in microservices versus monoliths.

Full core replacement

Replace the core banking system with a new platform, bought or built, and migrate all accounts and products onto it. This is the path with the largest potential benefit and the most public failures.

Approach Relative cost Delivery risk Typical time Removes the knowledge problem When it fits
Rehost Low to medium Low to medium Months to about two years No Cost pressure or data center exit; rules are stable
Refactor and code conversion Medium Medium, driven by testing One to three years per major estate Partly, only if followed by redesign Bounded applications with good tests
Strangler around the core Medium, spread over years Low per step, cumulative over time Several years, with value early Yes, gradually, as rules are re-specified Most banks that need agility and cannot afford a single failure
Full core replacement Highest Highest Several years, often overrun Yes, if the migration re-specifies rules Simpler product sets, strong vendor platform, or unavoidable constraint

Why Big-Bang Replacements Fail, and What Successful Programs Did Differently

Public failures follow a consistent pattern: the new platform is ready in a technical sense, and the migration moves all customers at once onto a system whose operational behavior has not been proven at full load.

TSB, 2018. TSB migrated its customers from a platform operated by its former parent to a new platform in a single weekend in April 2018. The platform experienced immediate failures and, according to the UK's Financial Conduct Authority, all of TSB's branches and a significant proportion of its 5.2 million customers were affected by the initial problems. In December 2022 the FCA and the Prudential Regulation Authority fined TSB a combined 48.65 million pounds for failing to organize and control the migration adequately and to manage the risks of its outsourcing arrangements, per the FCA's announcement. TSB told the BBC in 2019 that the incident had cost it 330 million pounds.

Postbank and Deutsche Bank, 2022 to 2023. Deutsche Bank's Unity program moved roughly 12 million Postbank customers onto a common platform with Deutsche Bank's own German customers in several waves from 2022. Although the migration was phased, the German regulator BaFin reported considerable disruption in customer service in 2023, including long processing times for account closures and savings repayments, and appointed a special representative to monitor remediation in September 2023, per BaFin. The lesson is that phasing the data move is not enough if back-office operations are not ready for the new system.

Commonwealth Bank of Australia, two programs. CBA's core replacement on SAP began in 2008 with an original estimate of 580 million Australian dollars and had reached 1.1 billion by the halfway point, as reported by iTnews in 2011. The bank migrated accounts in stages rather than in one conversion, and the program delivered a real-time core, at a cost far above the original estimate. In March 2026 CBA reported moving that SAP core onto cloud infrastructure after an 18-month project that ran five full dress rehearsals, two more than planned, and a final cutover through which, the bank says, customers could keep transacting.

What the successful programs share is visible in the contrast:

  • Incremental migration. Accounts, products or regions move in slices, each small enough that a failure affects a bounded population and can be reversed.
  • Parallel running. Old and new systems process the same transactions and results are reconciled daily before the old system is switched off. Differences are the most valuable output of the program, because each one is an undocumented rule surfacing.
  • Data before logic. Replicating ledger data to the new environment first, and proving it reconciles, lets new services read real data and exposes data-quality problems long before any logic moves.
  • Rehearsal at full scale. Cutovers are rehearsed end to end, repeatedly, including rollback, until the rehearsal is uneventful.
  • Operations readiness, not just technology readiness. Back-office processes, customer service and exception handling are proven on the new system before the customer population moves.

The broader patterns for sequencing legacy change are set out in legacy modernization for executives.

Where AI Code Tools Help, and Where They Are Oversold

The most credible public use of AI on legacy code so far is comprehension rather than conversion. Morgan Stanley said in 2025 that DevGen.AI, an internal tool launched that January, had reviewed nine million lines of legacy code and saved an estimated 280,000 developer hours by translating old code into plain-English specifications that engineers then used to rewrite it. The company's own figures are the only source for the savings, and the design choice is the notable part: the tool explains, engineers rewrite.

That aligns with where these tools are strongest:

  • Documentation and explanation. Summarizing what a program does, generating descriptions of data flows, and mapping which jobs touch which files.
  • Rule extraction. Proposing candidate business rules from code for human experts to confirm, which directly attacks the knowledge problem.
  • Dependency mapping. Identifying which programs, jobs and data sets are connected, so migration slices can be drawn safely.
  • Test generation. Proposing test cases from existing behavior, which is the most expensive part of any conversion.

They are oversold for safe automatic rewrite. A model that translates COBOL into Java can produce plausible code that differs from the original on an edge case nobody tested, which is the same class of error as a fabricated fact in any other generated output. For a ledger, a subtle difference in rounding or date handling is not a bug to fix later; it is a misstatement of customer balances. Automated conversion is viable only when paired with parallel running and reconciliation that would catch such a difference, at which point the tool is an accelerator for a disciplined program rather than a substitute for one.

A Decision Framework by Institution Type

Institution type Binding constraint Recommended path
Large universal bank with a heavily customized in-house core Risk of a single failure; volume of undocumented rules Strangler around the core, API layer first, data replication, AI-assisted documentation; keep the mainframe ledger until slices are proven
Regional bank on an ageing in-house core Change lead time blocking products Strangler for new products on a modern platform; migrate legacy products in parallel-run slices; consider full replacement only for a simple product set
Community bank or credit union on a vendor-hosted core Dependence on the vendor's roadmap and contract terms Modernize around the core through the vendor's APIs; negotiate contract terms before considering a switch
Bank after a merger, running two cores Duplicate cost and operational complexity Consolidate onto the stronger core in incremental waves with operations readiness gates, not a single conversion weekend
Institution under acute cost pressure, stable products Licence and hardware spend Negotiate pricing; offload reads; rehost if needed; do not start a replacement for cost reasons alone
Digital bank or new entrant None; no legacy core Choose a modern core, and design for the day it becomes legacy: documented rules, full test coverage, exportable data

The position, and its tradeoff

For most banks with a significant mainframe estate, the strangler pattern with data moved first is the right default, and the mainframe should be expected to remain the ledger of record for years while that happens. It turns one very large, very risky decision into many small reversible ones, it recovers the undocumented rules as a by-product of parallel running, and it delivers new capability to customers long before the core is retired.

The tradeoff is duration and double running. For several years the bank pays for both estates, maintains the integration between them, and reconciles their outputs. Programs that lose executive sponsorship midway can end up with the costs of both and the benefits of neither, a permanently half-strangled core. A full replacement is justified when the product set is simple enough to re-specify completely, when a strong vendor platform fits with limited customization, or when a constraint such as an unsupported platform removes the option to wait. Even then, the lesson of the public failures is to migrate in slices and rehearse until the cutover is boring.

Frequently Asked Questions

Why do banks still use mainframes?

Because mainframes process very high volumes of integrity-critical transactions reliably, and because the banking logic running on them encodes decades of business rules that are often documented nowhere else. Replacing the hardware is achievable. Recovering and re-proving every rule, product and edge case on a new system is what makes replacement slow, risky and expensive.

Is COBOL the main problem with banking mainframes?

No. COBOL is an old language but a learnable one. The harder problem is that many business rules exist only in the code, some written by people long gone, and nobody can confidently predict the effect of changing them. That knowledge gap, together with capacity-based pricing and a retiring skills pool, is the real cost of the estate.

How much does mainframe modernization cost?

It varies too widely by approach and estate for a single figure to be meaningful, and public examples show large overruns. Commonwealth Bank of Australia's SAP core replacement was originally estimated at 580 million Australian dollars and had reached 1.1 billion by its halfway point. Rehosting costs far less and changes less. The largest costs usually sit in testing, parallel running and reconciliation rather than in software or conversion.

Can AI convert COBOL to a modern language automatically?

AI tools can translate COBOL into modern languages, but safe automatic rewrite of core banking logic remains unproven, because a plausible translation can differ from the original on an untested edge case. The strongest current use is comprehension: documenting what programs do, extracting candidate business rules for experts to confirm, and generating tests. Any conversion needs parallel running and reconciliation before it touches customer balances.

What is the strangler pattern in banking modernization?

It is an approach where new capabilities are built outside the legacy core and functions are gradually routed away from it, product by product or process by process, until the remaining core is small enough to retire. It usually starts with an API layer in front of the mainframe and data replication to a modern store. It spreads risk across many small, reversible steps at the cost of running two estates for several years.

The Bottom Line

The mainframe is not the reason banks struggle to change. It is where decades of decisions about how the bank works have been stored in a form that only the machine can fully read. Replacing it without recovering those decisions moves the problem to a new platform, and doing so all at once turns every unrecovered rule into a potential customer-facing failure on the same weekend.

The programs that succeed accept that modernization is a multi-year migration of data, rules and operations, and they design for reversibility at every step. They move data first, run old and new side by side until the differences stop surprising anyone, use AI to read the estate rather than to rewrite it blind, and switch the mainframe off last. The cost of that discipline is time. The cost of skipping it is on public record.