What problem does it solve?
Broadcasters and publishers sit on archives that go back decades: newscasts, interviews, documentaries and raw, unpublished footage, described by whatever labelling convention was in use at the time it was catalogued, often no more than a title, a date and a one line description written by hand. Content makers who want to reuse this material, for an anniversary package, a breaking story that needs context, or a compilation, have to search on those sparse terms or scrub through hours of tape themselves.
Modern content operations make this worse before AI makes it better: the volume of new footage grows every year, on demand and streaming platforms want the archive searchable at the level of a clip rather than a programme, and content makers want to search for specific shots in natural language, such as the ABC's example of a cricketer with zinc on their nose, rather than the exact keywords a manual index was built around. Multimodal models can watch video and listen to audio, and describe both in natural language, which makes richer, clip level metadata practical at archive scale.
How does it work?
- Ingest. Video and audio files, old and new, are pulled into an AI enhanced digital asset or archive management system, either as a one off backfill of the historical archive or as part of the daily ingest pipeline for new footage.
- Analyse. A multimodal model segments the file and, for each segment, describes what is shown (people, objects, scenes, on screen text and logos), transcribes and translates what is said, and recognises known people and public figures, going well beyond fixed keyword lists to describe the actual visual and audio content.
- Structure the output. The generated descriptions are turned into structured, searchable metadata: entities, timecodes, categories and free text descriptions, consistent across decades of inconsistent source material.
- Human review. A documentalist or archivist validates and corrects the AI generated metadata, particularly for sensitive categories (identifying named individuals, historically significant events), and that correction feeds back into improving the model over time.
- Search and reuse. Content makers search the archive in natural language rather than exact keywords, and can retrieve a specific clip in seconds instead of scrubbing through the source tape, freeing time for editorial work rather than manual search.
- Audience
- Back office
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | 1 million | 1 | 1 vendor |
Value drivers: Employee productivity, Lower cost to serve, Revenue growth.
Indicative value
A broadcaster with a 50,000 hour video archive
USD 375,000 to USD 1.8 million
Archive cataloguing cost avoided one off
How this is calculated
Formula: archiveHours * humanTaggingCostPerHour * aiCoverageShare. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Archive hours to catalogue archiveHours, hours | 50,000 | 50,000 | The reference archive. |
| Fully loaded cost of manual cataloguing per hour of footage humanTaggingCostPerHour, USD per hour of footage | 15 | 40 | Editorial assumption for archivist or documentalist time. Replace with your own cost. |
| Share of the archive AI tagging can cover before human review aiCoverageShare, fraction of archive hours | 0.5 | 0.9 | Editorial assumption, replace with your own. Panorama Audiovisual reports RTVE's medium term goal of cataloguing 198,220 hours of material from its regional centres, and a 2025 contracted scope of 20,000 hours (whose printed components, 18,000 plus 5,000 plus 5,000, sum to 28,000); either figure is a small share of that goal. The ABC's one million video records analysed within two weeks says nothing about the share of an archive AI can cover. Neither source supports a specific coverage share. |
What it leaves out: Gross cataloguing cost avoided only. It leaves out the cost of running the AI service, the documentalist time still needed to validate and correct output, and the licensing revenue or editorial value the newly discoverable footage may unlock.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Radiotelevisión Española (RTVE)
Spain · Media and entertainment · 2026
NexTReT describes an AI based automatic metadata service for RTVE's Documentary Archive, integrated into its ARCA document management system with an interface for documentalists to validate the results, to make decades of audiovisual heritage material searchable and reusable. The case study is undated and does not say when the service ran or whether it later ended. Separately, RTVE ran a 2025 tender for automatic archive metadata, reported by Panorama Audiovisual; see the verification note for why that tender is not folded into this record.
No outcome disclosed.
Australian Broadcasting Corporation
Australia · Media and entertainment · 2025
The ABC modernized its 90 year old archive with Gemini in Vertex AI, adding rich, AI generated metadata to CoDA (Content Digital Archives), its open access platform for journalists, producers and editors. Gemini catalogs video segments, flags incidental footage, and powers semantic search so content makers can describe what they are looking for in natural language instead of scanning vague, human written tags. The editability of the AI generated metadata is presented as a key advantage, with human oversight reviewing and refining output for consistency and to meet the ABC's quality standards, and a human in the loop correction process feeding back into the models over time. The ABC has since integrated the tagging into its daily content pipeline so new uploads are described automatically.
- Interactions handled: 1 million, within two weeks
"Analysed one million video records within two weeks with the Gemini API in Vertex AI"
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Digitized video and audio files, or a digitization pipeline for analogue source material
- A target metadata schema (entities, categories, timecodes) that both legacy and AI generated records can populate
- Existing catalogue records to migrate or reconcile with newly generated metadata
Systems to integrate
- Digital asset or media asset management (MAM) system
- Multimodal AI model or vendor platform for video and audio analysis
- Search index that serves natural language queries over the generated metadata
- Documentalist or archivist review and correction interface
Complexity: Medium
The AI analysis itself is largely a managed multimodal model call; the harder work is the archive's own metadata schema and validation workflow, migrating inconsistent decades old records into a structure the AI's output can populate, and integrating with the existing asset management system rather than building a parallel one.
- 1
Pilot on a bounded, representative slice
Choose a few thousand hours that span the archive's range of formats and eras before committing to the full backfill, and measure both metadata quality and documentalist correction time.
- 2
Design the metadata schema before the pipeline
Agree the target categories, entities and timecodes the organization actually needs to search on, so the AI's output is structured to that schema from the start rather than reworked later.
- 3
Build the human review step in, not on
Give documentalists an interface to validate and correct AI generated metadata as part of the ingest workflow, especially for named people and sensitive events, rather than treating review as an afterthought.
- 4
Prioritise by reuse value, not just volume
Catalogue the material most likely to be reused first (anniversaries, recurring formats, frequently requested topics) so the pilot demonstrates value quickly, then expand to the full backfill.
- 5
Integrate into daily ingest
Once quality is proven on the backfill, apply the same pipeline to new footage as it is ingested, so the archive stops growing its backlog even as historical material catches up.
Guardrails
- Human validation of AI generated metadata before it is published as authoritative, particularly identification of named individuals
- A documented policy for what the system may not do, such as making rights or licensing decisions from metadata alone
- Access controls on sensitive archive material that mirror the organization's existing editorial and legal restrictions
- Version history so a correction to AI generated metadata is auditable and reversible
KPIs to instrument
- Search time for a content maker to locate a specific clip, before and after
- Share of the archive with AI generated metadata, backfill and new intake separately
- Documentalist correction rate on AI generated metadata, by category
- Archive material reused in new productions or licensed, before and after
Human in the loop
Documentalists and archivists validate and correct AI generated metadata as part of the ingest workflow, with particular attention to identifying named people, historically sensitive footage and anything that will inform a licensing or rights decision; their corrections are the mechanism that improves the model's output over time, not a one time quality check.
Common failure modes
- Confident but wrong identification
- Multimodal models can misidentify people or events with fluent, plausible sounding metadata. Require human validation before any AI identified person or event is treated as authoritative.
- Metadata schema drift
- AI generated categories that do not map cleanly to the archive's existing schema fragment search rather than improving it. Fix the target schema before the pipeline runs at scale.
- Backlog blindness
- A backfill project that does not also cover new intake just moves the backlog forward in time. Build the pipeline for daily ingest from the start, even if the backfill runs separately.
- Sensitive content surfaced without control
- Making decades of raw, unpublished footage newly searchable can surface material that was never meant for wide internal access. Apply the organization's existing access and editorial controls to AI generated search results, not just to the original files.
What are the risks and rules?
EU AI Act
Depends on design
Cataloguing objects, scenes, logos and spoken content is not listed in Annex III and is typically minimal risk. Annex III point 1(a) covers remote biometric identification: the automated, one to many matching of a person's face or voice, without their active involvement and typically at a distance, against a reference database of identified individuals to establish who they are, in so far as its use is permitted under relevant Union or national law; it excludes one to one biometric verification. A feature that recognises and names a specific person in archive footage by comparing them against such a database meets that definition and is high risk, while grouping similar looking footage without assigning an identity does not. Article 6(3) lets a provider assess a listed system itself as not high risk when it performs only a narrow procedural task, but that derogation is unlikely to cover a system whose purpose is naming an individual, so treat person recognition as high risk by default. Point 1(b) covers biometric categorisation, inferring a sensitive or protected attribute from a person's face or voice. RTVE's 2025 contract specifies speaker gender identification as a feature; whether that counts as a protected attribute under point 1(b) is contested, since Recital 54 ties that category to attributes protected under GDPR Article 9(1), and sex or gender is not listed there. Treat gender inference from voice as potentially high risk and apply the same governance the organization uses for other biometric systems until that question is settled.
Guidance
- Annex III: High-Risk AI Systems Referred to in Article 6(2) (European Union, Europe). Point 1(a) covers remote biometric identification, the automated, one to many comparison of a face or voice against a database of identified individuals, relevant where archive tagging recognises and names specific people. Point 1(b) covers biometric categorisation, inferring a sensitive attribute; whether gender inferred from voice, such as RTVE's 2025 contracted speaker gender identification feature, falls under this point is contested and treated here as potentially high risk rather than settled.
Controls to put in place
- Human validation of any AI generated identification of a named person before publication
- Inventory entry for the tagging system, its model provider and the categories of personal data it processes
- Documented retention and access policy for AI generated metadata, aligned with the archive's existing editorial controls
- Bias and accuracy checks on person recognition across demographic groups before wide rollout
Frequently asked questions
- Can AI tag an entire decades old video archive automatically?
- It can generate a first pass of rich metadata at a scale manual cataloguing cannot match. The ABC analysed one million video records within two weeks using Gemini in Vertex AI. RTVE has pursued automatic metadata for its archive since 2020 and, under a 2025 tender, contracted Crosspoint and Amplify to automatically analyse 20,000 hours of content; separately, a NexTReT case study (undated) describes a similar service built into RTVE's ARCA system, with an interface for documentalists to validate the results.
- Does AI metadata tagging replace archivists and documentalists?
- No, not in the deployments we found. The ABC describes a human in the loop correction process that reviews and refines the AI generated metadata to ensure consistency and meet its quality standards, and NexTReT's RTVE deployment built an interface for documentalists to validate the results. We recommend treating human review of named people and sensitive material as a guardrail on any such system, not an afterthought.
- Is recognising people in archive footage a high risk use under the EU AI Act?
- It can be. Object, scene and logo detection is typically minimal risk, but recognising a specific named person by matching their face or voice against a database of known individuals is remote biometric identification under Annex III point 1(a), which is high risk. RTVE's 2025 contract specifies speaker gender identification as a feature, not confirmed live by any source we found; whether inferring gender from voice counts as biometric categorisation under point 1(b) is contested, since that category is tied to attributes protected under GDPR and sex is not among them. We recommend treating gender inference from voice as potentially high risk and applying the same governance as other biometric systems.
- What does a broadcaster get back for cataloguing its archive with AI?
- Faster search for content makers (the ABC's case study describes finding the clip they need in seconds instead of taking an hour to find the right video and then scrubbing through hours of tape), a larger share of the archive that is actually discoverable and reusable, and, for organizations that license footage, a larger catalogue of material that can be found and cleared for reuse in the first place.
How to cite this page
Blits.ai AI Use Case Library, "AI metadata tagging and indexing for media archives", last verified 28 September 2026, https://www.blits.ai/ai-use-cases/media-archive-metadata-tagging. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 28 September 2026: First published