AI use case

AI for predictive network maintenance in telecom

Machine learning that spots the early signs of network failure, such as degrading cells, faulty customer equipment, ageing hardware or planned digging near fibre, and triggers a preventive fix, a remote reset or a targeted intervention before customers lose service.

By Len Debets · Last verified 27 September 2026 · 6 public deployments

At least 10 million
Interactions handled
Verizon (organization claim).
USD 600,000 to USD 12 million
Indicative value per year
A national fixed and mobile operator with about 40,000 customer affecting network faults a year. Worked example, see how it is calculated.

What problem does it solve?

Much network maintenance is reactive or calendar based. Faults are found when an alarm fires or when customers call, and field teams replace equipment on a schedule whether it needs it or not. Many failures give warning signs first: a cell whose throughput slowly degrades, a modem that keeps dropping its connection, a router with rising error counts, a battery that no longer holds its charge. Those signals sit in performance data that nobody has time to watch.

Some outages have nothing to do with the equipment itself. Verizon notes that every year thousands of fiber lines are damaged by accidental cuts during construction and excavation, which can affect customers' connectivity for anything from a few hours to several days. A fault that reaches the customer can cost a support call, a technician visit and some goodwill. The opportunity is to act on the warning signs early enough to fix the problem remotely, during a planned window, or before the digger arrives.

How does it work?

  1. Gather the signals. Performance counters, alarms, device telemetry from customer equipment, environmental and power data from sites, and external data such as dig requests or weather.
  2. Score the risk. Models learn the patterns that preceded past failures and score each cell, line, device or site for the probability of failure or degradation in the coming days.
  3. Classify the likely cause. For each at risk element the system proposes the probable root cause (hardware, configuration, interference, power, external damage) so the right fix is chosen.
  4. Act at the right level. Low risk fixes such as a remote reset, a configuration rollback or a customer equipment reboot run automatically within limits; hardware swaps and site visits are scheduled as planned work; external risks trigger outreach, such as contacting an excavator.
  5. Learn from the outcome. Every prevented and every missed failure is fed back to retrain the models and tune the thresholds.
Audience
Back office
Autonomy
Supervised agent
Adoption
Early adopters
Channels
Internal tools, API and system to system

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI for predictive network maintenance in telecom
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
2.5 million to 10 million
22 organization

Value drivers: Customer experience, Lower cost to serve, Risk and loss reduction, Speed and cycle time.

Indicative value

A national fixed and mobile operator with about 40,000 customer affecting network faults a year

USD 600,000 to USD 12 million

Fault handling cost avoided per year

How this is calculated

Formula: faults * preventableShare * costPerFault. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Customer affecting network faults per year faults, faults per year20,00060,000Editorial assumption for a national operator. Replace with your own fault volume.
Share of faults prevented or fixed before customers are affected preventableShare, fraction of faults0.10.25Editorial assumption, not calibrated by any source on this page. Telstra's 2.5 million SmartFix proactive actions in FY25 show the scale of such programs but say nothing about the share of faults that can be prevented; the share of your own faults that show warning signs is the number to measure first.
Cost of a customer affecting fault costPerFault, USD per fault300800Editorial assumption covering repair, technician visits and customer contacts. Replace with your own fully loaded cost.

What it leaves out: Direct fault cost only. It leaves out avoided service level penalties, churn and complaint handling, the extra cost of preventive work on false alarms, and the cost of the data platform.

Who already uses it?

6 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Telstra

Australia · Telecommunications · 2025

ScaledGrade B

Telstra's SmartFix system is embedded in its network operations and automatically fixes many issues before customers notice a problem. Telstra reports the number of proactive actions it performed in FY25 and says they prevented nearly 1 million support calls. The blog post is part of Telstra's description of its wider AI program and gives no detail on the models or the types of fixes.

  • Interactions handled: 2.5 million, FY25, proactive actions
    "It automatically fixes many issues before customers notice a problem – in FY25 it performed 2.5 million proactive actions, preventing nearly 1 million support calls by resolving issues in advance."
    Claimed by: organization

Orange

France · Telecommunications · 2024

ProductionGrade B

After a two year production trial on the French backbone, Orange Global Network and an SD-WAN network, Orange added the Augtera Network AI platform to its NOC tools. Topology based auto correlation groups alarms so operations experts see far fewer of them, and anomaly detection on metrics and logs flags weak signals so incidents can be handled before customers notice. Orange and Augtera say the correlation will cut the daily number of alarms the NOC has to address by 70%, presented as the expected effect of the rollout rather than a measured result. The integration was due to start in April 2024 in Orange Global Networks, an IP network with thousands of routers in 800 points of presence across 100 countries, with full rollout planned by the end of 2024.

No outcome disclosed.

Verizon

United States · Telecommunications · 2024

ProductionGrade B

Verizon uses artificial intelligence and machine learning on the 811 call before you dig requests it receives to identify the excavations most likely to damage its underground fiber. The model weighs historical and current activity at the location and the past record of the excavator on site, and high risk digs trigger preventive steps such as extra communication with the excavator. The solution is integrated with Verizon's 811 system; Verizon describes the potential benefit but has not published a measured reduction in fiber cuts.

  • Interactions handled: at least 10 million, per year, 811 dig requests screened
    "Verizon is utilizing advanced artificial intelligence (AI) and machine learning techniques to sort through over ten million 811 dig requests annually to identify high-risk excavations."
    Claimed by: organization

Telefónica España

Spain · Telecommunications · 2025

ProductionGrade C

Telefónica España built a network data platform on Microsoft Azure (Azure Data Explorer, Azure Databricks and Power BI) to store and analyse the large volumes of data its 4G and 5G mobile network produces. The team uses it for anomaly detection, to address issues before they affect customers, and for automated network optimization. Telefónica says the project is live with several use cases deployed and that results have been very positive; Microsoft's summary adds substantial savings in operating costs. No figures are published.

No outcome disclosed.

KDDI

Japan · Telecommunications · 2022

ScaledGrade C

KDDI deployed Nokia's AVA Performance Degradation Detection and Resolution (PDDR) solution nationwide to monitor its 4G and 5G radio network around the clock. The model detects performance degradations that raise no alarm, so called silent cells, classifies the likely root cause, and hands recoverable cases to KDDI's own recovery system, which tries to fix them automatically. Recovered cells feed back into the training data. KDDI started on 4G in 2019 and extended the system to its 5G NSA network in 2021.

No outcome disclosed.

Vodafone

United Kingdom · Telecommunications · 2021

ProductionGrade C

Vodafone and Nokia jointly developed an Anomaly Detection Service, based on Nokia Bell Labs technology and running on Google Cloud, that detects and troubleshoots irregularities such as mobile site congestion, interference and unexpected latency before they affect customers. After an initial deployment on more than 60,000 4G cells in Italy, it was being rolled out across Vodafone's European network in July 2021, with all European markets planned by early 2022 and plans to apply it later to 5G and core networks. Vodafone expected around 80 percent of its anomalous mobile network issues and capacity demands to be detected and addressed automatically; no measured result was published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Historical performance and alarm data per network element, kept for at least a year
  • Fault and repair records with root cause codes that can be joined to that data
  • Telemetry from customer premises equipment where fixed access is in scope
  • Site power, battery and environmental data
  • External data sources where relevant, such as dig request notifications

Systems to integrate

  • Performance and fault management systems
  • Network inventory and topology
  • Workforce management and field scheduling
  • Device management platforms for customer equipment
  • Trouble ticketing and customer notification systems

Complexity: High

Predicting failures needs years of clean history that links alarms and performance data to the faults that followed. Many operators have the data but not the labels, and acting on predictions means changing how field and NOC work is planned.

  1. 1

    Pick a failure mode with a clear payoff

    Start with one failure that is frequent, costly and preceded by measurable signals, such as degrading cells, unstable customer equipment or fibre damage from digging.

  2. 2

    Build the labelled history

    Join past faults to the data that preceded them. This is usually most of the work and the main reason projects stall.

  3. 3

    Decide the action before the model

    Agree what happens when the score is high: an automated reset, a planned visit or a call to a third party. A prediction nobody acts on is only a report.

  4. 4

    Automate only the safe fixes

    Let the system run reversible remote actions within limits and route everything else to engineers and field planners with the evidence attached.

  5. 5

    Measure prevented and missed failures

    Track both, because a model that raises many alarms can look busy while missing the failures that matter.

Guardrails

  • Automatic actions limited to reversible, low impact fixes with rollback and rate limits
  • Hardware swaps and site visits approved by planners, not triggered directly by a score
  • Maintenance windows and change freezes respected by every automated action
  • Human review of any action that affects many customers at once

KPIs to instrument

  • Faults prevented, measured against a comparable control group of elements
  • Precision of predictions, as the share of flagged elements that really degraded
  • Customer contacts and technician visits per thousand customers
  • Mean time between failures for the targeted element types
  • Automated actions that had to be rolled back

Human in the loop

Engineers set the thresholds and the list of automated fixes, planners approve preventive visits, and the NOC can pause automation at any time. A sample of automated actions is reviewed every week against what actually happened to the element afterwards.

Common failure modes

Predictions without actions
The model scores risk accurately but nobody owns the follow up. Tie every score band to a named action and owner.
Preventive work on healthy equipment
Too many false alarms send technicians to sites that were fine. Track precision and the cost of each preventive visit.
Automation that causes outages
An automated reset during peak hours takes down more customers than the fault would have. Use windows, rate limits and rollback.
Drift after network change
New equipment and software releases change what normal looks like. Retrain and revalidate after major upgrades.

What are the risks and rules?

EU AI Act

Depends on design

Scoring failure risk and planning maintenance is normally minimal risk. Annex III point 2 lists AI systems intended as safety components in the management and operation of critical digital infrastructure as high risk, and public electronic communications networks fall within that infrastructure. Recital 55 limits safety components to systems that directly protect the physical integrity of the infrastructure or the health and safety of persons and property, and excludes components used solely for cybersecurity. An operator whose automated actions meet that test must treat the system as high risk.

Guidance

Controls to put in place

  • Documented list of automated actions with owners, limits and rollback procedures
  • Change control for model thresholds and automation rules
  • Audit trail linking each prediction to the action taken and the outcome
  • Periodic validation of model precision and recall on recent failures

Frequently asked questions

What does predictive maintenance look like at scale in a telecom operator?
Telstra reports that its SmartFix system performed 2.5 million proactive actions in FY25, fixing many issues before customers noticed and preventing nearly 1 million support calls. Nokia reported in December 2022 that KDDI monitors its 4G and 5G radio network around the clock with Nokia's AVA PDDR, which detects silent cell degradations that raise no alarm and hands them to KDDI's automatic recovery system.
Is it only about network equipment?
No. Verizon applies machine learning to more than ten million 811 dig requests a year to identify high risk excavations near its fiber, and takes preventive steps such as extra communication with the excavator. The same approach can target customer premises equipment, site power and batteries.
Where should an operator start?
With one frequent, costly failure mode that has measurable warning signs and an agreed action, and with the labelled history that links past faults to the data that preceded them.

How to cite this page

Blits.ai AI Use Case Library, "AI for predictive network maintenance in telecom", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/predictive-network-maintenance. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Telecommunications

AI copilot for network operations centre fault triage

AI in the network operations centre (NOC) that correlates alarms and performance data from radio, transport, core and fixed networks into a small number of probable faults, ranks them by customer impact, proposes the likely root cause and fix from runbooks, vendor documentation and past tickets, and routes the ticket to the right team, while an engineer decides what to change.

Deployments
6 public, best grade B
Reported cycle time reduction
at least 95%
Deutsche Telekom, organization claim
Telecommunications

AI copilot for field technicians and dispatch optimization

AI that decides whether a fault needs a site visit at all, predicts what work and parts a job will need, helps plan and update appointments, and gives technicians on site guided diagnosis and instant answers from manuals and past jobs, so more jobs are fixed on the first visit.

Deployments
2 public, best grade B
Autonomy
Copilot
Telecommunications

Agentic AI for autonomous, intent based network operations

AI agents that run closed loops over a telecom network: they take an intent from the operator (for example a latency or availability target for a service), observe the network, diagnose deviations and execute corrective actions across radio, transport and core, within guardrails set by engineers and with human approval for major changes.

Deployments
6 public, best grade B
Reported cycle time reduction
at least 95%
Deutsche Telekom, organization claim
Telecommunications

AI agent for network outage detection and customer communication

An AI agent that turns network alarms into a clear picture of which customers are affected by an outage and why, tells them proactively by message, app or phone with a cause and an estimated fix time, answers their questions during the incident, and updates them until service is restored.

Deployments
1 public, best grade B
Autonomy
Supervised agent
Telecommunications

AI agent for device and connectivity troubleshooting on voice and chat

An AI agent that diagnoses and fixes a customer's broadband, mobile, TV or device problem by conversation on the phone or in chat, running line tests and remote resets through the operator's systems, guiding the customer step by step, and booking an engineer or handing over to a technician when the fault needs a person.

Deployments
4 public, best grade B
Median containment rate
70%
3 deployments