AI use case

AI transcription, subtitles and captions for audio and video

AI that transcribes recorded audio and video, such as podcasts, broadcasts, lessons, interviews and hearings, in several languages, separates the speakers and produces timed transcripts, subtitles and captions for a human editor to check, delivered as files for publishing or the archive.

By Len Debets · Last verified 26 September 2026 · 5 public deployments

50%
Reported cost reduction
Warner Bros. Discovery, vendor claim.
80%
Reported cycle time reduction
Warner Bros. Discovery, vendor claim.
USD 225,000 to USD 1.5 million
Indicative value per year
A regional broadcaster or publisher that captions 5,000 hours of content a year. Worked example, see how it is calculated.

What problem does it solve?

Every organization that publishes audio or video has the same backlog: captions for people who are deaf or hard of hearing, subtitles for viewers who speak another language, and searchable transcripts for the archive. Done by hand, captioning one hour of content takes many hours of transcription, timing and translation. Ateme describes up to 15 hours of manual work per hour of video per language, and external providers that were slow and hard to scale. So organizations caption the flagship content and leave the rest, or subtitle into one language only.

The gap is growing as content volumes grow and accessibility rules tighten. SVT, Sweden's public broadcaster, says it simply could not caption its local news without automation, because it publishes for 21 regional stations several times a day. Speech recognition is now accurate enough that the human role shifts from typing to editing: the machine produces a timed draft with speakers marked, and an editor fixes names, terms and meaning before publication.

  • The World Health Organization estimates that over 5% of the world's population, or 430 million people, require rehabilitation for disabling hearing loss.Deafness and hearing loss (2026)

How does it work?

  1. Ingest the file. A new recording arrives from the media asset system, learning platform, podcast host or court recording system, and a job starts automatically.
  2. Transcribe with timings. Speech recognition produces text with word level timestamps and detects the spoken language, using a custom vocabulary of names and terms. Pacers Sports & Entertainment reduced its transcription error rate by 87% by tuning the model to its own broadcasts.
  3. Separate speakers. Diarization marks who spoke when, so transcripts read as a dialogue and captions can show speaker changes.
  4. Translate and segment. The text is translated into the target languages and cut into subtitle lines that respect reading speed, line length and shot changes.
  5. Edit and approve. An editor reviews the draft in a subtitle editor, corrects names, numbers and meaning, and approves each language before release.
  6. Deliver files. The system exports standard formats such as SRT and WebVTT captions and a plain transcript, and returns them to the publishing platform and the archive.

Live captioning, as on arena screens or live broadcasts, uses the same speech models but removes the editor, so accuracy tuning and filters matter even more.

Audience
Back office
Autonomy
Copilot
Adoption
Mainstream
Channels
API and system to system, Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI transcription, subtitles and captions for audio and video
KPIMedianReported rangeData pointsClaimed by
Cost reductionToo few to pool
50%
11 vendor
Cycle time reductionToo few to pool
80%
11 vendor
Error reductionToo few to pool
87%
11 vendor

Value drivers: Lower cost to serve, Speed and cycle time, Inclusion and access, Compliance quality, Employee productivity.

Indicative value

A regional broadcaster or publisher that captions 5,000 hours of content a year

USD 225,000 to USD 1.5 million

Captioning cost avoided per year

How this is calculated

Formula: hours * costPerHour * costReduction. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Hours of content captioned per year hours, hours of content per year5,0005,000The reference organization.
Current cost of captioning one hour of content costPerHour, USD per hour of content150600Editorial assumption covering in house or outsourced transcription, timing and review in one language. Replace with your own rates.
Share of captioning cost removed costReduction, fraction of cost0.30.5Conservative against the benchmark on this page (Google Cloud reports a 50% reduction in overall costs for Warner Bros. Discovery's AI captioning tool), because editing time remains.

What it leaves out: Covers the existing captioning volume only. It leaves out the cost of running the speech models and the editing tool, the value of content that is captioned or subtitled for the first time, extra languages, and the audience and compliance benefits.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Ateme

France · Media and entertainment · 2025

ProductionGrade C

Ateme, a French video compression and delivery company, added a subtitle step to its file transcoding platform: after transcoding, Gemini on Vertex AI transcribes the audio, spots timecodes and generates subtitles in the requested languages, and a script converts them to SRT for the client's workflow. Ateme says a job that took up to 15 hours of manual work per hour of video now takes minutes and costs less than a dollar per hour of content.

No outcome disclosed.

Pacers Sports & Entertainment

United States · Media and entertainment · 2025

ProductionGrade C

Pacers Sports & Entertainment trained a custom speech model on hundreds of hours of its own game broadcasts, with lists of player, coach and official names, to caption live announcers for fans who are deaf, hard of hearing or do not speak English, on arena screens and in its mobile apps. English came first, then Spanish, and Microsoft reports the team is adding 12 more languages. It is the live variant of the job: captions are delivered in real time, and Microsoft reports built in moderation filters that help keep inappropriate or misinterpreted content off the screen. The system now captions Pacers, Fever and All Star games at Gainbridge Fieldhouse.

  • Error reduction: 87%
    "By tailoring the model to the Pacers’ broadcast style, they reduced the speech-to-text transcription error rate by 87%—a level of precision that made it possible to extend captioning from mobile apps to arena screens with confidence."
    Claimed by: vendor

Comeen

France · Technology and software · 2024

ProductionGrade C

Comeen, a workplace software company, launched automatic multilingual subtitles for videos shown on its clients' digital signage at the end of 2024. Employees upload a video in their usual tools and receive subtitles in 40 languages, produced by a multi step AI workflow that the company says makes them usable as is. It replaces a process that used several providers over several days.

No outcome disclosed.

Warner Bros. Discovery

United States · Media and entertainment · 2024

ProductionGrade C

Warner Bros. Discovery built an AI captioning tool on Vertex AI. Google Cloud reports that it delivered a 50% reduction in overall costs and cut the time to caption a file by 80% compared with manual captioning. No detail on languages, volumes or the review step is published.

  • Cost reduction: 50%
    "Warner Bros. Discovery built an AI captioning tool with Vertex AI, delivering a 50% reduction in overall costs and an 80% reduction in the time it takes to manually caption a file without the use of machine learning."
    Claimed by: vendor
  • Cycle time reduction: 80%
    "Warner Bros. Discovery built an AI captioning tool with Vertex AI, delivering a 50% reduction in overall costs and an 80% reduction in the time it takes to manually caption a file without the use of machine learning."
    Claimed by: vendor

Sveriges Television (SVT)

Sweden · Media and entertainment · 2021

ScaledGrade C

SVT, Sweden's public broadcaster, transcribes its video content and generates closed captions automatically with Azure speech services, in production since 2021. The broadcaster says it could not caption its local news otherwise, because it publishes for 21 regional stations at the same time several times a day, and that feedback from viewers with hearing loss is mostly positive.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Access to the source audio or video files, in a quality good enough for speech recognition
  • A glossary of names, places, brands and technical terms per programme, course or court
  • House style for captions and subtitles (line length, reading speed, speaker labels, sound descriptions)
  • A sample of manually captioned content to measure accuracy before and after

Systems to integrate

  • Media asset management system, learning platform, podcast host or recording system
  • Subtitle or transcript editor for human review
  • Publishing platform or video player that accepts SRT or WebVTT files
  • Archive or search index for transcripts

Complexity: Low

Speech recognition, diarization and machine translation are mature and available from many providers. The work is in the workflow around them: getting files in and out of the media or learning platform, a glossary of names and terms, an editing step that editors actually like, and output formats that the players accept.

  1. 1

    Measure your baseline first

    Take a sample of content across programme types, speakers and languages, caption it the current way and record the time and cost. Without this baseline, no one can say whether the AI draft saves time once editing is counted.

  2. 2

    Compare speech engines on your own audio

    Run the same sample through several speech recognition engines and compare word error rate, especially on names, accents, dialects and overlapping speech. Results on vendor benchmarks rarely match results on your material.

  3. 3

    Build the glossary and tune

    Feed names and domain terms to the engine as custom vocabulary or a correction step. Pacers Sports & Entertainment tuned its model with its own broadcasts and name lists and pushed the error rate far below its original target.

  4. 4

    Put the editor at the centre

    Give editors a tool that shows the draft with timings, speaker labels and low confidence words highlighted, and measure editing time per hour of content. The editor, not the model, signs off what is published.

  5. 5

    Add languages one at a time

    Start with same language captions, then add subtitle languages where audience data shows demand, each with a native speaker review of a sample before going live.

  6. 6

    Automate delivery

    Trigger jobs when files arrive and return approved SRT, WebVTT and transcript files to the publishing platform and archive automatically, so nothing depends on manual uploads.

Guardrails

  • A human editor approves every caption and subtitle file before publication
  • Custom vocabulary and a correction step for names, places and terms
  • Low confidence words and segments flagged for the editor, not silently published
  • Profanity and sensitive word checks on output, especially for content aimed at children
  • Recordings and transcripts of identifiable people processed and stored under the organization's retention rules

KPIs to instrument

  • Word error rate per programme type, language and speaker profile, on a monthly sample
  • Editing time per hour of content, compared with the manual baseline
  • Cost per hour of captioned content, including model and editing costs
  • Share of published content with captions and with subtitles, per language
  • Errors reported by viewers or learners after publication

Human in the loop

Editors review and correct every file before it is published, with extra attention to names, numbers, quotes and anything legally sensitive. For translated subtitles, a native speaker checks a sample per language and programme type. For live captioning, where no editor can intervene, a producer monitors the output and can switch captions off.

Common failure modes

Offensive or embarrassing misrecognition
The engine hears an innocent word as an offensive one, and it reaches the screen. Prevent with sensitive word checks and editor review, above all for children's content.
Names and terms mangled
People, places and technical terms are transcribed wrongly, which undermines trust and can be defamatory. Maintain a glossary per programme and check names first in review.
Review that becomes a rubber stamp
Editors under time pressure approve drafts without reading them. Track editing time and sample published files for errors.
Captions that do not fit the screen
Correct text that is badly timed or too long to read. Enforce reading speed and line length rules in the segmentation step.

What are the risks and rules?

EU AI Act

Depends on design

Transcription and captioning are not listed in Annex III and are not a prohibited practice under Article 5, so the tier depends on how captions are published. Article 50(4) requires deployers to disclose AI generated or manipulated text published to inform the public on matters of public interest, such as news captions, unless it has undergone human review or editorial control and someone holds editorial responsibility, so the editor step keeps most deployments outside this duty. The provider duty to mark output in Article 50(2) does not apply where the system does not substantially alter the input or its semantics, which fits same language transcription better than translated subtitles. Unreviewed news captions or subtitles should therefore be disclosed as automatic.

Guidance

  • Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). Paragraph 4 sets the disclosure duty for AI generated text published to inform the public on matters of public interest and exempts content under human review or editorial control; paragraph 2 exempts systems that do not substantially alter the input or its semantics.
  • Guidelines 02/2021 on virtual voice assistants (European Data Protection Board, Europe). Final version of July 2021 on processing voice data in voice assistants; its analysis of voice recordings as personal data, and when they become biometric data, also applies to recordings and transcripts of identifiable speakers.

Controls to put in place

  • Documented editorial sign off per published file
  • Glossary and custom vocabulary with an owner per programme or course
  • Retention and access rules for recordings and transcripts of identifiable people
  • Monthly accuracy sample with word error rate reported per language
  • Labelling of automatic captions where no human review takes place

When it went wrong elsewhere

Frequently asked questions

How much does AI captioning save?
Published results are large but come from vendors. Google Cloud reports that Warner Bros. Discovery's AI captioning tool delivered a 50% reduction in overall costs and cut the time to caption a file by 80%. Your saving depends on how much editing your content needs, so measure editing time per hour on a sample first.
Can AI captions be published without a human check?
For live events there is often no human in the loop before the caption appears, and organizations such as Pacers Sports & Entertainment rely on tuned models and moderation filters. For recorded content, an editor should check names, numbers and sensitive words, and under the EU AI Act human editorial review also removes the duty to label published text as AI generated.
How is this different from AI meeting notes?
Meeting tools produce summaries and action items for participants. This use case produces a complete, timed transcript and caption files for publication or the archive, where every word and its timing matter and an editor signs off.
Which accessibility rules apply?
In the EU, the European Accessibility Act requires services that give access to audiovisual media, such as players and apps, to transmit subtitles for the deaf and hard of hearing with adequate quality and in sync with sound and video. Automation is how broadcasters such as SVT caption content they could not caption by hand.

How to cite this page

Blits.ai AI Use Case Library, "AI transcription, subtitles and captions for audio and video", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/audio-and-video-transcription-and-captioning. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryGovernment and public sector

AI meeting summarization and action items

AI that summarizes internal and operational meetings, such as team, project, board and case meetings: it transcribes an online or in person meeting with the participants' knowledge and produces a summary, decisions and action items with owners and dates for the organizer to check and share. It is the general purpose tool; client advice meetings and sales calls, which feed a regulated record or a sales pipeline, have their own pages.

Deployments
5 public, best grade B
Autonomy
Copilot
Government and public sector

AI translation and interpretation for multilingual public services

AI that translates government content, documents and conversations between officials and the public, in writing and in real time speech, so people can use public services in their own language, with human translators and interpreters reviewing what carries legal or safety weight.

Deployments
8 public, best grade B
Autonomy
Copilot
Government and public sector

AI for court and case file summarization

AI that condenses court filings, case files, evidence recordings and earlier decisions into structured summaries, chronologies and draft case reports with references to the source pages, so that judges, prosecutors, tribunal staff and government lawyers find what matters faster, while the person responsible reads the underlying material and makes every legal judgment.

Deployments
4 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI quality and compliance monitoring of every customer interaction

Automated quality assurance that transcribes and scores every customer interaction, voice and chat, against the organization's own rubric, checking required disclosures and script adherence, flagging conduct and mis selling risk, and surfacing coaching opportunities, instead of the small sample a human QA team can review.

Deployments
5 public, best grade C
Reported quality score uplift
about 10%
British Gas, vendor claim