What problem does it solve?
Core systems in many large organizations still run on code written decades ago, often in COBOL or proprietary languages, with little documentation. The evidence on this page shows both: a global bank whose foundational systems were designed decades ago, at a time when COBOL talent is becoming scarce, and Toyota Motor Europe, whose applications in a proprietary language depend on developers who are retiring. Every change is slow and risky, and a full rewrite is hard to plan because nobody can say with confidence what the old system actually does.
The hard part of modernization was never typing the new code. It is comprehension (what does this program do, which rules are buried in it, what depends on it) and proof (does the new version behave the same on real data). Three of the four deployments on this page use AI for comprehension: Morgan Stanley and Toyota Motor Europe turn code into readable specifications and documentation, and at the GFT bank AI also generated test scenarios and made the converted code more readable, while deterministic tools did the conversion. Amazon went furthest, using a code transformation agent to help migrate applications to a newer Java version.
How does it work?
- Inventory and dependency mapping. Tools parse the estate to find programs, copybooks, jobs, data stores and the calls between them, so work can be split into modules.
- Explain the code. A model generates technical and business documentation per program: what it does, its inputs and outputs, and the business rules and conditions it applies.
- Review by the remaining experts. Subject matter experts check a sample of the documentation against the code and correct it. Their corrections improve the next batch.
- Convert or rewrite. Deterministic converters or engineers produce the modern code from the specification, with AI assistance for readability and idiomatic structure.
- Prove equivalence. AI generates regression tests and test data from the documented rules; old and new systems run in parallel on production like data until the differences are explained.
- Cut over in controlled steps. Each module moves through the normal change process, with traceability from the legacy module to the new service.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Cost savings | Not pooled | USD 260 million | 1 | 1 organization |
| Cycle time reduction | Too few to pool | about 50% | 1 | 1 organization |
| Hours saved | Not pooled | about 280,000 hours | 1 | 1 organization |
Value drivers: Speed and cycle time, Lower cost to serve, Risk and loss reduction, Employee productivity.
Indicative value
A bank with 5 million lines of legacy code in scope for modernization
USD 1.2 million to USD 6 million
Analysis and test design effort avoided over the program
How this is calculated
Formula: linesOfCode / 1000 * hoursPerThousandLines * aiReduction * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Lines of legacy code in scope linesOfCode, lines of code | 5,000,000 | 5,000,000 | The reference organization. |
| Analysis, documentation and test design hours per thousand lines, done manually hoursPerThousandLines, hours per thousand lines | 10 | 20 | Editorial assumption. Replace with the estimate from your own modernization plan. |
| Share of that effort the AI removes aiReduction, fraction of effort | 0.3 | 0.5 | Editorial assumption, conservative against the evidence on this page (Morgan Stanley reports roughly 280,000 hours saved on nine million lines, about 31 hours per thousand lines). |
| Blended cost of an engineering hour hourlyCost, USD per hour | 80 | 120 | Editorial assumption. Replace with your own rates, including specialist contractors. |
What it leaves out: Covers comprehension, documentation and test design only. It leaves out conversion, parallel running, infrastructure and license savings after decommissioning, and the risk reduction of having documented systems, which is often the larger benefit.
Who already uses it?
5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
United States · Technology and software · 2026
Google used an internal LLM based system to help engineers with large scale code migrations, such as changing identifier types from int32 to int64. Engineers doing the migrations estimated the total time spent was reduced by about 50%, and reported that 80% of the code changes in landed changelists were AI authored, with the rest written by humans.
- Cycle time reduction: about 50%
"The total time spent on the migration was reduced by an estimated 50% as reported by the engineers doing the migration."
Claimed by: organization
Airbnb
United States · Technology and software · 2025
Airbnb migrated around 3,500 React test files from Enzyme to React Testing Library using a combination of frontier language models and automation. Airbnb had originally estimated the migration would take about 1.5 years of engineering time to do by hand, and instead completed it in 6 weeks.
No outcome disclosed.
Amazon
United States · Technology and software · 2024
Amazon integrated the Java transformation capability of Amazon Q Developer into its internal systems and migrated tens of thousands of production applications from Java 8 or 11 to Java 17 with its assistance. AWS describes the product's transformation agents as analyzing source code, generating new code, testing it and executing the change once the customer approves; the sources do not describe how Amazon's own developers reviewed each upgrade. Amazon estimates more than 4,500 years of development work saved compared with manual upgrades, and annual savings from hosts it could remove after the faster Java 17 runtime. AWS has since extended the approach to .NET, VMware and mainframe workloads.
- Cost savings: USD 260 million, per year, estimated from hosts removed after the Java 17 upgrade
"This effort saved more than 4,500 years of development work, compared to what it would have taken previously, and realized performance improvements of $260 million in annual cost savings."
Claimed by: organization
Toyota Motor Europe
Belgium · Automotive · 2026
Toyota Motor Europe runs more than 70 custom applications on legacy mainframe and AS400 platforms, several written in a proprietary language with little documentation and a shrinking pool of experts. With Deloitte and the AWS Generative AI Innovation Center it built a proof of concept on Amazon Bedrock that generates technical documentation, business documentation and process flows from the source code of a warranty handling application of over 1.3 million lines. The remaining experts reviewed a sample against the source and confirmed its accuracy. The proof of concept covered 2 of the 10 modules; AWS reports that it has since led to a production rollout.
No outcome disclosed.
Morgan Stanley
United States · Capital markets · 2025
Morgan Stanley launched DevGen.AI in January 2025, an in house tool built on OpenAI's GPT models and trained on the languages in its own code base, including company specific ones. It turns code in older languages such as Perl into plain English specifications that developers then use to rewrite the code in modern languages. The firm keeps developers in the loop because the tool does not yet write the new code as well as a human, and said it would not cut its engineering workforce as a result.
- Hours saved: about 280,000 hours, first five months after launch
"Mike Pizzi, Morgan Stanley’s global head of technology and operations, told WSJ that in the five months since its launch, DevGen.AI has worked through nine million lines of code, saving the firm’s 15,000 developers roughly 280,000 hours of work."
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Complete source code, including copybooks, job control and configuration
- Access to the remaining experts for review of generated documentation
- Production like test data, masked where it contains personal data
- An agreed target architecture and coding standards
Systems to integrate
- Source repositories and mainframe code management
- Static analysis and dependency mapping tools
- Model access in a tenant with zero retention and suitable data location
- Test automation and parallel run comparison tooling
Complexity: High
The model is rarely the hardest part. Modernization programs touch the most critical systems, need parallel running and, in regulated firms, early engagement with supervisors, and depend on experts who are scarce. AI shortens comprehension and testing; it does not remove the program.
- 1
Start with comprehension, not conversion
Pick one module and have the AI document it. Let the experts grade the documentation. This tells you quickly how reliable the model is on your code and languages.
- 2
Split the estate into modules
Use dependency mapping to find boundaries where a module can move on its own, and sequence the program by business risk and dependency.
- 3
Choose the conversion approach per module
Deterministic conversion keeps logic identical but produces unidiomatic code; rewriting from a specification gives better code but more risk. Many programs combine both.
- 4
Generate and run the tests
Turn the documented rules into regression tests and compare old and new outputs on the same data. Differences are investigated, not waved through.
- 5
Keep traceability
Link every new service back to the legacy programs and documented rules it replaces, for auditors and for the next change.
- 6
Retire the old code
Plan decommissioning from the start. Savings only arrive when the legacy runtime is switched off.
Guardrails
- No direct cutover from AI output; every module passes testing and parallel running
- Expert review of generated documentation before it is used as a specification
- Code processed only in a tenant with zero retention and agreed data location
- Traceability from each legacy module to its replacement
- Change approval by the owners of the business process, not only by IT
KPIs to instrument
- Documentation accuracy on expert reviewed samples
- Hours per module for analysis and test design, before and after
- Differences found in parallel running and their root causes
- Defects after cutover per module
- Legacy capacity decommissioned
Human in the loop
Experts validate the documentation, engineers own the new code, and business owners sign off on behavioral equivalence after parallel running. Human sign off at the cutover gate is not optional.
Common failure modes
- Plausible but wrong documentation
- The model describes what similar code usually does, not what this code does. Expert sampling and generated tests against real behavior catch it.
- Converting dead code
- Large parts of old estates are unused. Measure what runs before converting everything.
- Losing the business rules
- Rules hidden in data or job control are missed when only programs are analyzed. Include the whole runtime in scope.
- Big bang cutover
- A full switch without parallel running turns small differences into incidents. Move module by module.
What are the risks and rules?
EU AI Act
Minimal risk
Tools that analyze, document and translate code are not prohibited practices under Article 5 and are not listed in Annex III, so no high risk obligations apply to the tooling. Engineers and analysts know they are working with an AI tool, including when they query the documentation through a chat assistant, so the Article 50 disclosure duty has no practical effect for the deploying organization. What remains is AI literacy for the staff who use it (Article 4). If the system being modernized is itself an AI system in an Annex III area (for example creditworthiness assessment, point 5(b)), its new version still has to meet the high risk requirements.
Rules that apply
Guidance
- Guidelines on Risk Management Practices, Technology Risk (Monetary Authority of Singapore, Asia Pacific). Supervisory expectations for technology risk governance, system development, testing and change management at financial institutions in Singapore.
- Guidelines for secure AI system development (UK National Cyber Security Centre, Europe). Security guidance for organizations that build AI systems, relevant to in house pipelines and agents that process and generate code.
Controls to put in place
- Program level risk assessment with the AI tooling in scope
- Third party risk review of model providers and integrators
- Test evidence and parallel run results retained per module
- Traceability records from legacy to new components
- Independent review of cutover readiness for critical systems
Frequently asked questions
- Can AI convert COBOL to Java on its own?
- Not reliably enough for core systems. Most deployments on this page use AI for comprehension and documentation, and humans or deterministic tools produce the new code. Morgan Stanley's DevGen.AI turns legacy code into English specifications that developers rewrite, because the firm says the tool does not yet write new code as well as a human, and GFT describes a bank where deterministic tools did the conversion while generative AI produced documentation and test scenarios, the only deployment here that reports AI generated tests. Amazon's agent did help upgrade applications from Java 8 or 11 to Java 17, a version upgrade rather than a change of language.
- How much time does it save?
- Morgan Stanley says DevGen.AI worked through nine million lines in five months and saved about 280,000 developer hours. Amazon reports that its code transformation agent helped migrate tens of thousands of production applications to Java 17 and estimates that this saved more than 4,500 years of development work. Toyota Motor Europe's documentation proof of concept gives no time figure, and savings on mainframe estates depend heavily on how much expert review the output needs.
- What should stay with humans?
- Validation of business rules, the decision to cut over, and sign off on test and parallel run results. The retiring experts are most valuable as reviewers of AI generated documentation.
How to cite this page
Blits.ai AI Use Case Library, "AI for legacy code modernization", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/legacy-code-modernization. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published