What problem does it solve?
Every organization that publishes audio or video has the same backlog: captions for people who are deaf or hard of hearing, subtitles for viewers who speak another language, and searchable transcripts for the archive. Done by hand, captioning one hour of content takes many hours of transcription, timing and translation. Ateme describes up to 15 hours of manual work per hour of video per language, and external providers that were slow and hard to scale. So organizations caption the flagship content and leave the rest, or subtitle into one language only.
The gap is growing as content volumes grow and accessibility rules tighten. SVT, Sweden's public broadcaster, says it simply could not caption its local news without automation, because it publishes for 21 regional stations several times a day. Speech recognition is now accurate enough that the human role shifts from typing to editing: the machine produces a timed draft with speakers marked, and an editor fixes names, terms and meaning before publication.
- The World Health Organization estimates that over 5% of the world's population, or 430 million people, require rehabilitation for disabling hearing loss.Deafness and hearing loss (2026)
How does it work?
- Ingest the file. A new recording arrives from the media asset system, learning platform, podcast host or court recording system, and a job starts automatically.
- Transcribe with timings. Speech recognition produces text with word level timestamps and detects the spoken language, using a custom vocabulary of names and terms. Pacers Sports & Entertainment reduced its transcription error rate by 87% by tuning the model to its own broadcasts.
- Separate speakers. Diarization marks who spoke when, so transcripts read as a dialogue and captions can show speaker changes.
- Translate and segment. The text is translated into the target languages and cut into subtitle lines that respect reading speed, line length and shot changes.
- Edit and approve. An editor reviews the draft in a subtitle editor, corrects names, numbers and meaning, and approves each language before release.
- Deliver files. The system exports standard formats such as SRT and WebVTT captions and a plain transcript, and returns them to the publishing platform and the archive.
Live captioning, as on arena screens or live broadcasts, uses the same speech models but removes the editor, so accuracy tuning and filters matter even more.
- Audience
- Back office
- Autonomy
- Copilot
- Adoption
- Mainstream
- Channels
- API and system to system, Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Cost reduction | Too few to pool | 50% | 1 | 1 vendor |
| Cycle time reduction | Too few to pool | 80% | 1 | 1 vendor |
| Error reduction | Too few to pool | 87% | 1 | 1 vendor |
Value drivers: Lower cost to serve, Speed and cycle time, Inclusion and access, Compliance quality, Employee productivity.
Indicative value
A regional broadcaster or publisher that captions 5,000 hours of content a year
USD 225,000 to USD 1.5 million
Captioning cost avoided per year
How this is calculated
Formula: hours * costPerHour * costReduction. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Hours of content captioned per year hours, hours of content per year | 5,000 | 5,000 | The reference organization. |
| Current cost of captioning one hour of content costPerHour, USD per hour of content | 150 | 600 | Editorial assumption covering in house or outsourced transcription, timing and review in one language. Replace with your own rates. |
| Share of captioning cost removed costReduction, fraction of cost | 0.3 | 0.5 | Conservative against the benchmark on this page (Google Cloud reports a 50% reduction in overall costs for Warner Bros. Discovery's AI captioning tool), because editing time remains. |
What it leaves out: Covers the existing captioning volume only. It leaves out the cost of running the speech models and the editing tool, the value of content that is captioned or subtitled for the first time, extra languages, and the audience and compliance benefits.
Who already uses it?
5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Ateme
France · Media and entertainment · 2025
Ateme, a French video compression and delivery company, added a subtitle step to its file transcoding platform: after transcoding, Gemini on Vertex AI transcribes the audio, spots timecodes and generates subtitles in the requested languages, and a script converts them to SRT for the client's workflow. Ateme says a job that took up to 15 hours of manual work per hour of video now takes minutes and costs less than a dollar per hour of content.
No outcome disclosed.
Pacers Sports & Entertainment
United States · Media and entertainment · 2025
Pacers Sports & Entertainment trained a custom speech model on hundreds of hours of its own game broadcasts, with lists of player, coach and official names, to caption live announcers for fans who are deaf, hard of hearing or do not speak English, on arena screens and in its mobile apps. English came first, then Spanish, and Microsoft reports the team is adding 12 more languages. It is the live variant of the job: captions are delivered in real time, and Microsoft reports built in moderation filters that help keep inappropriate or misinterpreted content off the screen. The system now captions Pacers, Fever and All Star games at Gainbridge Fieldhouse.
- Error reduction: 87%
"By tailoring the model to the Pacers’ broadcast style, they reduced the speech-to-text transcription error rate by 87%—a level of precision that made it possible to extend captioning from mobile apps to arena screens with confidence."
Claimed by: vendor
Comeen
France · Technology and software · 2024
Comeen, a workplace software company, launched automatic multilingual subtitles for videos shown on its clients' digital signage at the end of 2024. Employees upload a video in their usual tools and receive subtitles in 40 languages, produced by a multi step AI workflow that the company says makes them usable as is. It replaces a process that used several providers over several days.
No outcome disclosed.
Warner Bros. Discovery
United States · Media and entertainment · 2024
Warner Bros. Discovery built an AI captioning tool on Vertex AI. Google Cloud reports that it delivered a 50% reduction in overall costs and cut the time to caption a file by 80% compared with manual captioning. No detail on languages, volumes or the review step is published.
- Cost reduction: 50%
"Warner Bros. Discovery built an AI captioning tool with Vertex AI, delivering a 50% reduction in overall costs and an 80% reduction in the time it takes to manually caption a file without the use of machine learning."
Claimed by: vendor - Cycle time reduction: 80%
"Warner Bros. Discovery built an AI captioning tool with Vertex AI, delivering a 50% reduction in overall costs and an 80% reduction in the time it takes to manually caption a file without the use of machine learning."
Claimed by: vendor
Sveriges Television (SVT)
Sweden · Media and entertainment · 2021
SVT, Sweden's public broadcaster, transcribes its video content and generates closed captions automatically with Azure speech services, in production since 2021. The broadcaster says it could not caption its local news otherwise, because it publishes for 21 regional stations at the same time several times a day, and that feedback from viewers with hearing loss is mostly positive.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Access to the source audio or video files, in a quality good enough for speech recognition
- A glossary of names, places, brands and technical terms per programme, course or court
- House style for captions and subtitles (line length, reading speed, speaker labels, sound descriptions)
- A sample of manually captioned content to measure accuracy before and after
Systems to integrate
- Media asset management system, learning platform, podcast host or recording system
- Subtitle or transcript editor for human review
- Publishing platform or video player that accepts SRT or WebVTT files
- Archive or search index for transcripts
Complexity: Low
Speech recognition, diarization and machine translation are mature and available from many providers. The work is in the workflow around them: getting files in and out of the media or learning platform, a glossary of names and terms, an editing step that editors actually like, and output formats that the players accept.
- 1
Measure your baseline first
Take a sample of content across programme types, speakers and languages, caption it the current way and record the time and cost. Without this baseline, no one can say whether the AI draft saves time once editing is counted.
- 2
Compare speech engines on your own audio
Run the same sample through several speech recognition engines and compare word error rate, especially on names, accents, dialects and overlapping speech. Results on vendor benchmarks rarely match results on your material.
- 3
Build the glossary and tune
Feed names and domain terms to the engine as custom vocabulary or a correction step. Pacers Sports & Entertainment tuned its model with its own broadcasts and name lists and pushed the error rate far below its original target.
- 4
Put the editor at the centre
Give editors a tool that shows the draft with timings, speaker labels and low confidence words highlighted, and measure editing time per hour of content. The editor, not the model, signs off what is published.
- 5
Add languages one at a time
Start with same language captions, then add subtitle languages where audience data shows demand, each with a native speaker review of a sample before going live.
- 6
Automate delivery
Trigger jobs when files arrive and return approved SRT, WebVTT and transcript files to the publishing platform and archive automatically, so nothing depends on manual uploads.
Guardrails
- A human editor approves every caption and subtitle file before publication
- Custom vocabulary and a correction step for names, places and terms
- Low confidence words and segments flagged for the editor, not silently published
- Profanity and sensitive word checks on output, especially for content aimed at children
- Recordings and transcripts of identifiable people processed and stored under the organization's retention rules
KPIs to instrument
- Word error rate per programme type, language and speaker profile, on a monthly sample
- Editing time per hour of content, compared with the manual baseline
- Cost per hour of captioned content, including model and editing costs
- Share of published content with captions and with subtitles, per language
- Errors reported by viewers or learners after publication
Human in the loop
Editors review and correct every file before it is published, with extra attention to names, numbers, quotes and anything legally sensitive. For translated subtitles, a native speaker checks a sample per language and programme type. For live captioning, where no editor can intervene, a producer monitors the output and can switch captions off.
Common failure modes
- Offensive or embarrassing misrecognition
- The engine hears an innocent word as an offensive one, and it reaches the screen. Prevent with sensitive word checks and editor review, above all for children's content.
- Names and terms mangled
- People, places and technical terms are transcribed wrongly, which undermines trust and can be defamatory. Maintain a glossary per programme and check names first in review.
- Review that becomes a rubber stamp
- Editors under time pressure approve drafts without reading them. Track editing time and sample published files for errors.
- Captions that do not fit the screen
- Correct text that is badly timed or too long to read. Enforce reading speed and line length rules in the segmentation step.
What are the risks and rules?
EU AI Act
Depends on design
Transcription and captioning are not listed in Annex III and are not a prohibited practice under Article 5, so the tier depends on how captions are published. Article 50(4) requires deployers to disclose AI generated or manipulated text published to inform the public on matters of public interest, such as news captions, unless it has undergone human review or editorial control and someone holds editorial responsibility, so the editor step keeps most deployments outside this duty. The provider duty to mark output in Article 50(2) does not apply where the system does not substantially alter the input or its semantics, which fits same language transcription better than translated subtitles. Unreviewed news captions or subtitles should therefore be disclosed as automatic.
Rules that apply
Guidance
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). Paragraph 4 sets the disclosure duty for AI generated text published to inform the public on matters of public interest and exempts content under human review or editorial control; paragraph 2 exempts systems that do not substantially alter the input or its semantics.
- Guidelines 02/2021 on virtual voice assistants (European Data Protection Board, Europe). Final version of July 2021 on processing voice data in voice assistants; its analysis of voice recordings as personal data, and when they become biometric data, also applies to recordings and transcripts of identifiable speakers.
Controls to put in place
- Documented editorial sign off per published file
- Glossary and custom vocabulary with an owner per programme or course
- Retention and access rules for recordings and transcripts of identifiable people
- Monthly accuracy sample with word error rate reported per language
- Labelling of automatic captions where no human review takes place
When it went wrong elsewhere
- YouTube's Captions Insert Explicit Language in Kids' Videos. Researchers found that automatic captions on videos from top children's channels contained inappropriate words the speakers never said, a risk for any unreviewed captioning.
Frequently asked questions
- How much does AI captioning save?
- Published results are large but come from vendors. Google Cloud reports that Warner Bros. Discovery's AI captioning tool delivered a 50% reduction in overall costs and cut the time to caption a file by 80%. Your saving depends on how much editing your content needs, so measure editing time per hour on a sample first.
- Can AI captions be published without a human check?
- For live events there is often no human in the loop before the caption appears, and organizations such as Pacers Sports & Entertainment rely on tuned models and moderation filters. For recorded content, an editor should check names, numbers and sensitive words, and under the EU AI Act human editorial review also removes the duty to label published text as AI generated.
- How is this different from AI meeting notes?
- Meeting tools produce summaries and action items for participants. This use case produces a complete, timed transcript and caption files for publication or the archive, where every word and its timing matter and an editor signs off.
- Which accessibility rules apply?
- In the EU, the European Accessibility Act requires services that give access to audiovisual media, such as players and apps, to transmit subtitles for the deaf and hard of hearing with adequate quality and in sync with sound and video. Automation is how broadcasters such as SVT caption content they could not caption by hand.
How to cite this page
Blits.ai AI Use Case Library, "AI transcription, subtitles and captions for audio and video", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/audio-and-video-transcription-and-captioning. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published