What problem does it solve?
Operator networks are large and multi vendor: Vodafone's first anomaly detection deployment alone covered more than 60,000 4G cells in Italy, and Orange Global Networks has thousands of IP routers in 800 points of presence. Each element raises its own alarms, so one broken fibre or a failed power supply produces a storm of alarms across radio, transport and core at the same time. NOC engineers spend the first part of every incident working out which alarms belong together, which domain owns the problem and whether customers are affected at all, while the ticket queue keeps growing.
The knowledge that shortens a fault (vendor manuals, runbooks, the fix that worked last month) is spread over many tools that do not talk to each other. Telstra describes the multi vendor problem directly: disparate vendor platforms cannot talk to each other, which slows fault diagnosis. Some degradations raise no alarm at all: Nokia, describing its work with KDDI, calls these "silent cells", which do not hurt service quality at once but eventually degrade the end user experience. The result is long mean time to repair, engineers who chase noise, and customers who report faults before the operator sees them.
How does it work?
- Collect and normalise. Alarms, performance counters, logs, topology and inventory from every vendor platform stream into one data layer, with the network topology as a graph.
- Correlate. Models group alarms that share a topological or temporal cause into one probable incident, so the NOC sees one problem instead of hundreds of alarms. Anomaly detection adds degradations that raised no alarm.
- Rank by customer impact. The incident is scored with traffic, affected services and customers, so a small fault on a busy site outranks a large one on an idle site.
- Explain and propose. A language model assistant summarises the evidence, retrieves matching runbook steps, vendor documentation and similar past tickets, and proposes the likely root cause and the next diagnostic or corrective step, with its sources.
- Route and track. The ticket goes to the owning team (radio, transport, core, field) with the correlated evidence attached; the engineer approves any change, and the outcome is fed back to improve correlation and ranking.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Cycle time reduction | Too few to pool | at least 95% | 1 | 1 organization |
| Interactions handled | Not pooled | at least 100 | 1 | 1 organization |
Value drivers: Employee productivity, Speed and cycle time, Customer experience, Lower cost to serve.
Indicative value
A mobile and fixed operator with a 24 hour NOC handling 100,000 network incidents a year
USD 300,000 to USD 5.4 million
NOC engineering time released per year
How this is calculated
Formula: incidents * triageMinutes / 60 * timeSaved * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Network incidents and trouble tickets triaged by the NOC per year incidents, incidents per year | 60,000 | 150,000 | Editorial assumption for a national operator. Replace with your own ticket volume. |
| Engineer minutes spent on correlation and diagnosis per incident triageMinutes, minutes per incident | 30 | 60 | Editorial assumption. Replace with a time study of your own NOC. |
| Share of triage time saved by correlation and assisted diagnosis timeSaved, fraction of triage time | 0.2 | 0.4 | Editorial assumption. Orange expects topology based correlation to cut the daily alarms its NOC has to address by 70%, but fewer alarms do not translate one to one into less engineer time, so the range stays well below that. |
| Fully loaded cost of a NOC engineer hour hourlyCost, USD per hour | 50 | 90 | Editorial assumption. Replace with your own fully loaded cost. |
What it leaves out: Engineering time only. It leaves out the larger effect of shorter outages on customer experience, churn and service level penalties, the contacts avoided when faults are fixed before customers call, and the cost of the data platform and integration work.
Who already uses it?
6 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Telstra
Australia · Telecommunications · 2026
In a proof of concept in a live telco cloud environment, Telstra showed an agentic AI capability that detected an unplanned infrastructure outage and resolved it autonomously by moving critical network applications to healthy hardware in minutes rather than hours. Using the Model Context Protocol with retrieval augmented generation, AI agents connect data from multiple vendor platforms into one view and give teams context aware recommendations to speed up fault resolution. Telstra presents it as groundwork for self healing, self optimising operations, not yet as a production service.
No outcome disclosed.
Bell Canada
Canada · Telecommunications · 2025
Bell deployed a network AI Ops solution built on Google Cloud that correlates network data with customer experience to rank issues by their real impact, so a small fault on a busy site gets attention before a larger one on a quiet site. Custom machine learning models handle detection and prioritisation, a graph of network relationships estimates customer impact, and Gemini models support incident analysis, retrieval of historical context and access to vendor documentation. Bell says the approach has significantly improved mean time to resolution but does not give a figure specific to AI Ops.
No outcome disclosed.
Deutsche Telekom
Germany · Telecommunications · 2025
Deutsche Telekom's RAN Guardian Agent, built with Gemini models on Google Cloud, went live in its German mobile network in November 2025. It is a multi agent system: one agent finds upcoming public events from public sources, another assesses whether nearby cells can carry the expected traffic and monitors them live, and a third executes corrective actions such as reallocating resources or adjusting configuration, documenting every action. It is being extended to the Czech Republic and Croatia. In February 2026 Deutsche Telekom announced MINDR, which applies the same approach end to end across radio, transport and core domains, with first production releases planned for later in 2026.
- Cycle time reduction: at least 95%, live operations, major events
"And in live operations it has reduced the time needed to manage major events from hours to around a minute, a more than 95% improvement."
Claimed by: organization - Interactions handled: at least 100, first month after launch, Christmas market events
"Since its launch in November 2025, RAN Guardian Agent has autonomously triggered over 100 remediation actions at Christmas market events during its first month."
Claimed by: organization
Orange
France · Telecommunications · 2024
After a two year production trial on the French backbone, Orange Global Network and an SD-WAN network, Orange added the Augtera Network AI platform to its NOC tools. Topology based auto correlation groups alarms so operations experts see far fewer of them, and anomaly detection on metrics and logs flags weak signals so incidents can be handled before customers notice. Orange and Augtera say the correlation will cut the daily number of alarms the NOC has to address by 70%, presented as the expected effect of the rollout rather than a measured result. The integration was due to start in April 2024 in Orange Global Networks, an IP network with thousands of routers in 800 points of presence across 100 countries, with full rollout planned by the end of 2024.
No outcome disclosed.
KDDI
Japan · Telecommunications · 2022
KDDI deployed Nokia's AVA Performance Degradation Detection and Resolution (PDDR) solution nationwide to monitor its 4G and 5G radio network around the clock. The model detects performance degradations that raise no alarm, so called silent cells, classifies the likely root cause, and hands recoverable cases to KDDI's own recovery system, which tries to fix them automatically. Recovered cells feed back into the training data. KDDI started on 4G in 2019 and extended the system to its 5G NSA network in 2021.
No outcome disclosed.
Vodafone
United Kingdom · Telecommunications · 2021
Vodafone and Nokia jointly developed an Anomaly Detection Service, based on Nokia Bell Labs technology and running on Google Cloud, that detects and troubleshoots irregularities such as mobile site congestion, interference and unexpected latency before they affect customers. After an initial deployment on more than 60,000 4G cells in Italy, it was being rolled out across Vodafone's European network in July 2021, with all European markets planned by early 2022 and plans to apply it later to 5G and core networks. Vodafone expected around 80 percent of its anomalous mobile network issues and capacity demands to be detected and addressed automatically; no measured result was published.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Alarm and event streams from every vendor element manager and OSS
- Performance counters per cell, link and node at a useful granularity
- An accurate network topology and inventory, ideally as a graph
- Historical trouble tickets with root cause and resolution codes
- Runbooks and vendor documentation in a searchable form
Systems to integrate
- Fault management and OSS platforms per domain (radio, transport, core, fixed access)
- Performance management and network data lake
- Network inventory and topology systems
- Trouble ticketing and workforce management
- Collaboration tools used by the NOC for incident communication
Complexity: High
The model is the easy part. The work is in getting clean, timely alarm, performance, topology and inventory data out of many vendor systems, and in earning the trust of engineers who have seen many correlation tools promise more than they delivered.
- 1
Start with one domain and one pain
Pick the domain with the worst alarm noise (often radio access) and measure today's alarms per real incident, time to diagnose and repeat tickets, so you have a baseline.
- 2
Fix the data before the model
Get topology and inventory right and stream alarms in near real time. Correlation built on a wrong topology produces confident nonsense.
- 3
Run in shadow mode
Let the system group alarms and propose root causes next to the existing process for several weeks, and have senior engineers grade each proposal before anyone relies on it.
- 4
Ground the assistant in your own knowledge
Load runbooks, vendor manuals and resolved tickets into retrieval with owners and review dates, and make the assistant cite its sources and say when it does not know.
- 5
Put suppression under change control
Every rule or model that hides alarms from engineers is a risk. Version it, test it against past major incidents and review it like any network change.
- 6
Widen to cross domain incidents
Once radio works, add transport and core so the system can trace a customer impact to its real cause across domains.
Guardrails
- The copilot proposes; engineers approve every configuration change and every ticket closure
- Suppressed alarms remain visible on request and are sampled for review every week
- Answers cite the runbook, document or past ticket they come from, with a refusal when nothing matches
- Read only access to network elements for the assistant, with any write actions routed through existing change management
- Major incidents always trigger the normal human escalation path, whatever the model says
KPIs to instrument
- Alarms per actionable incident, before and after correlation
- Mean time to detect and mean time to repair per domain
- Share of root cause proposals accepted by engineers
- Incidents first reported by customers rather than detected by the NOC
- Missed incidents, where a suppressed or low ranked alarm turned out to matter
Human in the loop
NOC engineers own diagnosis and every change to the network. The copilot correlates, ranks and drafts; engineers accept, correct or reject its proposals, and their corrections are the training signal. Senior engineers review suppression rules and a sample of closed incidents every week.
Common failure modes
- Suppression that hides a real outage
- A correlation rule groups a new failure under an old pattern and it never reaches an engineer. Sample suppressed alarms and replay past major incidents on every model change.
- Stale topology
- Correlation depends on knowing what is connected to what. When inventory lags behind the network, root cause proposals point at the wrong element.
- Fluent but wrong diagnosis
- A language model explains a fault convincingly from an outdated runbook. Require citations and track acceptance rates per runbook.
- Tool nobody opens
- The copilot lives in yet another screen. Put its output inside the ticket and the NOC wall, not beside them.
What are the risks and rules?
EU AI Act
Depends on design
The main test is Annex III point 2, which lists AI systems intended as safety components in the management and operation of critical digital infrastructure as high risk. A copilot that prepares diagnoses for engineers who decide every change is normally not such a safety component, and is then minimal risk. The tier rises when the system is designed to protect the safe operation of the network, for example by acting on it automatically to prevent or contain outages. Article 6(3) can exempt an Annex III system that only performs a preparatory task to an assessment and poses no significant risk of harm, provided the provider documents that assessment and registers the system.
Guidance
- Regulation (EU) 2024/1689 (AI Act), Annex III high risk AI systems (European Union, Europe). Point 2 covers AI systems intended as safety components in the management and operation of critical digital infrastructure.
- Directive (EU) 2022/2555 (NIS2 Directive) (European Union, Europe). Under Article 2(2), providers of public electronic communications networks and services fall under NIS2 regardless of their size, with risk management and significant incident reporting duties that the NOC tooling supports and must not weaken.
Controls to put in place
- Inventory entry for the copilot with an accountable owner and documented data sources
- Change control and regression tests on correlation and suppression rules
- Audit trail of every proposal, the engineer's decision and the outcome
- Weekly sampling of suppressed alarms and closed incidents by senior engineers
Frequently asked questions
- How much alarm noise can AI correlation remove?
- It depends on the network and the quality of topology data. After a two year production trial, Orange said topology based auto correlation from Augtera will cut the daily alarms its NOC has to address by 70%, stated as the expected effect of the rollout rather than a measured result. Measure your own alarms per real incident before and after, rather than relying on that figure.
- Does a generative AI copilot replace the correlation engine?
- No. Correlation and anomaly detection remain statistical and topology based. The language model sits on top: it explains the incident, finds the relevant runbook and past tickets, and drafts the next step. Bell's setup follows this pattern: custom AI and machine learning models correlate network data with customer experience to prioritise issues, and Gemini models support incident analysis, historical context retrieval and access to vendor documentation.
- Is a NOC copilot high risk under the EU AI Act?
- Usually not while engineers make every change. It can become high risk under Annex III point 2 when the system itself acts as a safety component in operating critical digital infrastructure, so the design choice about autonomy decides the tier.
How to cite this page
Blits.ai AI Use Case Library, "AI copilot for network operations centre fault triage", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/network-fault-triage-copilot. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published