AI use case

AI coding assistant for software developers

An AI assistant in the developer's IDE and code review flow that completes and generates code, explains unfamiliar modules, drafts unit tests and reviews pull requests for common defects, while generated code goes through the same review, testing and change controls as any other code.

By Len Debets · Last verified 27 September 2026 · 6 public deployments

20%
Median productivity gain
3 deployments.
At least 10.5 hours
Hours saved
CME Group (vendor claim).
USD 1.5 million to USD 7.5 million
Indicative value per year
An engineering organization with 500 developers. Worked example, see how it is calculated.

What problem does it solve?

Large organizations run thousands of engineers, and a big share of their time goes to work that is necessary but not differentiating: boilerplate, glue code, tests, reading code someone else wrote, and first pass review. Banks and insurers add a long tail of internal frameworks and legacy services that new joiners have to learn before they are productive.

Coding assistants take part of that load. The question for an engineering leader is no longer whether developers will use them, but how to capture the gain safely: keeping proprietary code out of external training, stopping insecure or unlicensed code from reaching production, and measuring real delivery rather than lines of code accepted.

How does it work?

  1. Inline completion and chat in the IDE. The assistant suggests code as the developer types and answers questions about the code base, using the open files and repository as context.
  2. Tests and explanations. Developers ask it to draft unit tests, explain a module or propose a fix for a failing build.
  3. Review assistance. On a pull request it summarizes the change and flags likely defects, which the human reviewer accepts or rejects.
  4. Normal controls apply. Generated code goes through peer review, static analysis, secret and licence scanning, tests and change approval, exactly like human code.
  5. Agentic tasks, carefully. Newer tools can plan and apply multi file changes or run commands. These should run in sandboxes with limited permissions, never directly against production.
Audience
Employee facing
Autonomy
Copilot
Adoption
Mainstream
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI coding assistant for software developers
KPIMedianReported rangeData pointsClaimed by
Productivity gain20%
8.7% to 42.4%
32 organization, 1 vendor
Hours savedNot pooled
at least 10.5 hours
11 vendor

Value drivers: Employee productivity, Speed and cycle time.

Indicative value

An engineering organization with 500 developers

USD 1.5 million to USD 7.5 million

Developer capacity released per year

How this is calculated

Formula: developers * loadedCost * affectedShare * timeSaved. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Developers with the assistant developers, developers500500The reference organization.
Fully loaded cost per developer loadedCost, USD per developer per year100,000150,000Editorial assumption. Replace with your own blended cost, including contractors.
Share of developer time spent on tasks the assistant helps with affectedShare, fraction of working time0.30.5Editorial assumption. Coding, tests and reading code, excluding meetings, design and incidents.
Time saved on those tasks timeSaved, fraction of task time0.10.2Editorial assumption, set below the Bank of America and ANZ figures on this page (Bank of America reports efficiency gains of over 20%; ANZ's controlled experiment measured about 42% less time on algorithmic Python challenges). Most CME Group developers using Gemini Code Assist report at least 10.5 hours a month, which falls inside this range. The only field randomized trial, Accenture's, measured a different thing (8.69% more pull requests), so replace this with your own control group result.

What it leaves out: Released capacity, not cash: it only becomes value if the time goes into more delivery. It leaves out licence and review costs, the extra review effort generated code can create, and quality effects in either direction.

Who already uses it?

6 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Bank of America

United States · Banking · 2025

ProductionGrade B

Bank of America software developers use a generative AI tool that helps them write and optimize code. The bank reported the efficiency gain in its April 2025 update on how its workforce uses AI, alongside its internal assistants and contact centre tools. The bank does not name the underlying model or vendor.

  • Productivity gain: at least 20%
    "Coding assistance – Bank of America software developers are using a GenAI-based tool to assist with code writing and optimization, through which they have experienced efficiency gains of over 20%."
    Claimed by: organization

ANZ

Australia · Banking · 2024

ProductionGrade B

ANZ, which employs over 5,000 engineers, ran a six week experiment with GitHub Copilot from June to July 2023: after two weeks of preparation, over 100 participants were split at random into a control group and a Copilot group that solved the same algorithmic Python challenges. The Copilot group took markedly less time and produced code with fewer code smells and bugs, while the effect on security was inconclusive. By the time of writing about 1,000 engineers were using Copilot. The study was written by ANZ staff and published on arXiv.

  • Productivity gain: 42.4%, two week A/B test on algorithmic Python challenges, self reported time
    "This study shows that Copilot improves the productivity of ANZ engineers by 42.36% on an average."
    Claimed by: organization

Citigroup

United States · Banking · 2024

ScaledGrade B

On Citi's fourth quarter 2024 earnings call, CEO Jane Fraser said the bank had armed 30,000 developers with AI tools to write code and launched two AI platforms for 143,000 colleagues. CIO Dive reported the same rollout as generative AI coding tools. No productivity or quality figures for the coding tools are disclosed, and the vendors are not named.

No outcome disclosed.

Meta

United States · Technology and software · 2024

ProductionGrade B

Meta built TestGen-LLM, a tool that uses large language models to draft and improve unit tests. During Instagram and Facebook test events, it improved 11.5% of all classes it was applied to, and 73% of its recommended test improvements were accepted by Meta software engineers for production deployment.

  • Accuracy: 73%
    "During Meta's Instagram and Facebook test-a-thons, it improved 11.5% of all classes to which it was applied, with 73% of its recommendations being accepted for production deployment by Meta software engineers."
    Claimed by: organization

CME Group

United States · Capital markets · 2025

ProductionGrade C

CME Group, which operates the Chicago Mercantile Exchange, gave its developers Gemini Code Assist. Google Cloud reports that most developers using it say they gain at least 10.5 hours a month.

  • Hours saved: at least 10.5 hours, per developer per month, self reported
    "CME Group, which operates the Chicago Mercantile Exchange, says most developers using Gemini Code Assist report a productivity gain of at least 10.5 hours a month."
    Claimed by: vendor

Accenture

Ireland · Professional services · 2024

ProductionGrade C

GitHub and Accenture ran a randomized controlled trial in which Accenture developers were randomly given GitHub Copilot or not, and measured DevOps telemetry such as pull requests and build success. The Copilot group opened more pull requests and saw many more successful builds, and surveyed developers reported less mental effort on repetitive tasks. GitHub also ran a company wide adoption analysis of installation and suggestion acceptance at Accenture.

  • Productivity gain: 8.7%
    "Ultimately, an increase in pull requests represents an increase in value delivered, and Accenture developers saw an 8.69% increase in pull requests."
    Claimed by: vendor

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Repository access rules and a list of repositories excluded from assistant context
  • Secure coding standards and approved libraries the assistant should follow
  • Baseline delivery metrics (cycle time, pull request throughput, change failure rate)

Systems to integrate

  • IDEs and the source control platform
  • CI pipeline with static analysis, secret scanning and licence scanning
  • Identity provider for licence assignment and single sign on
  • Model gateway or vendor tenant with zero retention terms

Complexity: Low

Rolling out a commercial assistant is technically simple. The effort is in the contract and data terms, security review, secure configuration of repositories and secrets, training, and a measurement plan that looks at delivery rather than acceptance rates.

  1. 1

    Settle data terms first

    Choose a tenant where your code is not retained or used for training, confirm where prompts are processed, and record the tool as a third party service in your risk register.

  2. 2

    Pilot with a control group

    Give the tool to a representative set of teams and compare them with similar teams without it over several weeks, on delivery metrics rather than surveys alone.

  3. 3

    Strengthen the pipeline

    Make secret scanning, dependency and licence checks and static analysis mandatory gates before merge, since more code will arrive faster.

  4. 4

    Train reviewers as well as authors

    Teach engineers to treat suggestions as untrusted input, to check generated tests actually test something, and to reject code they do not understand.

  5. 5

    Scale and keep measuring

    Roll out by team, track adoption and delivery metrics per cohort, and revisit settings when new agentic features arrive.

Guardrails

  • Zero retention and no training on the organization's code, confirmed contractually
  • Generated code passes the same review, testing and change approval as human code
  • Mandatory secret, dependency, licence and static analysis scanning before merge
  • Agentic features run in sandboxes without production credentials
  • Sensitive repositories excluded from assistant context where required

KPIs to instrument

  • Pull request throughput and lead time per team, compared with a control group
  • Change failure rate and escaped defects
  • Security findings per thousand lines in generated versus human code
  • Weekly active users among licensed developers
  • Developer satisfaction, surveyed quarterly

Human in the loop

A developer accepts or rejects every suggestion and remains the author of record. A second engineer reviews every change before merge, and release managers approve production changes as before. Agent generated multi file changes are reviewed like a new colleague's first pull request.

Common failure modes

Measuring the wrong thing
Acceptance rates and lines generated rise while delivery does not. Measure throughput, lead time and quality against a control group.
Faster insecure code
Suggestions reproduce insecure patterns or hard coded secrets. Scanning gates and reviewer training catch them; the tool alone does not.
Source code leakage
Engineers paste proprietary code into public chat tools when the approved tool is weak. Provide a good approved tool and block the alternatives.
Agents with too much reach
An agent with shell or database access acts outside its task. Limit permissions and never give it production credentials.

What are the risks and rules?

EU AI Act

Minimal risk

A coding assistant used by developers is not a prohibited practice under Article 5 and is not listed in Annex III. Developers know they are working with an AI tool, so the Article 50 disclosure duty has no practical effect for the deploying organization, and the marking of generated content under Article 50(2) falls on the tool's provider. What remains is AI literacy (Article 4). Using an AI system to monitor or evaluate individual developers' performance would fall under Annex III point 4(b), and the software the assistant helps build may itself fall under the Act.

Guidance

  • SP 800-218, Secure Software Development Framework (SSDF) Version 1.1 (NIST, North America). Baseline secure software development practices, such as code review, testing and vulnerability response, that apply to generated code as much as to code written by hand.
  • Guidelines for secure AI system development (UK National Cyber Security Centre, Europe). Guidelines for providers of AI systems in four areas (secure design, development, deployment, and operation and maintenance), relevant when you build your own tooling or agents around the assistant.
  • Article 4, AI literacy (European Union, Europe). Providers and deployers must take measures on the AI literacy of staff who use AI systems (the amended wording shown on this page asks them to support it rather than ensure a sufficient level). Here that means training developers on the tool's limits.

Controls to put in place

  • Third party risk assessment and register entry for the assistant vendor
  • Contractual zero retention and data location terms
  • Mandatory scanning gates in the CI pipeline
  • Developer training on reviewing generated code
  • Periodic review of agent permissions and repository exclusions

When it went wrong elsewhere

Frequently asked questions

How much faster do developers get with a coding assistant?
Published results vary widely with how they are measured. Bank of America reports efficiency gains of over 20% for its developers, Accenture's randomized trial with GitHub saw 8.69% more pull requests, and ANZ measured about 42% less time on algorithmic Python challenges. Set exercises tend to overstate everyday gains, so measure your own teams against a control group.
Does our code train the vendor's model?
It depends on the vendor, the plan and the contract, and terms differ between them. Confirm in writing that your code is not used for training and how long prompts are retained, check where prompts are processed, and treat the tool as a material third party service where your regulator expects that.
Is generated code a regulatory problem for a bank?
Not in itself. Rules such as DORA require ICT change management, testing and security controls, and these apply whoever or whatever wrote the code. The risk is the pipeline receiving more code than it can review, so strengthen scanning and review before scaling.

How to cite this page

Blits.ai AI Use Case Library, "AI coding assistant for software developers", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/developer-coding-assistant. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

AI for legacy code modernization

AI that reads legacy code such as COBOL, PL/I or old Java, explains what each program does, maps its data flows and dependencies, drafts the equivalent modern code or specification, and generates the regression tests needed to prove the new system behaves like the old one.

Deployments
5 public, best grade B
Reported cycle time reduction
about 50%
Google, organization claim
Cross industryBanking

AI assistant for developers integrating a company's APIs

An AI assistant on a developer portal and in its documentation that answers integration questions, recommends the right endpoints, helps debug connections and generates sample calls, grounded in the API catalogue, reference docs and test material, so clients and partners integrate faster with fewer support tickets.

Deployments
4 public, best grade B
Reported contact deflection
30%
Mapbox, organization claim
Cross industryBanking

AI for IT incident triage and root cause analysis (AIOps)

AI that turns a flood of monitoring alerts into one probable incident, routes it to the right team, proposes likely root causes and remediation from runbooks and past incidents, and drafts the stakeholder updates and the post incident review, while an engineer authorizes every change.

Deployments
5 public, best grade B
Median accuracy
90%
3 deployments
Cross industryBanking

Governed text to SQL analytics assistant

An assistant that turns a business user's plain language question into a query against governed data, runs it under that user's own data permissions and returns the table or chart together with the SQL and the tables used, so routine ad hoc questions no longer queue for the data team.

Deployments
3 public, best grade B
Autonomy
Assist
Cross industryBanking

AI enterprise knowledge search for employees

An assistant that lets any employee ask a question in plain language and get a synthesized answer from the organization's own policies, procedures, product manuals and research, with citations to the source documents and only from documents the employee is allowed to see.

Deployments
4 public, best grade B
Autonomy
Assist