What problem does it solve?
Platforms that let people post, chat or stream in real time cannot review that volume by hand. A live video feed produces a new frame to check every fraction of a second, and a chat or comment section can carry harassment, scams, CSAM, grooming or self harm content that needs to be caught in seconds, not after a user reports it. Purely human moderation teams cannot keep pace with this volume, and the delay between a violation appearing and a person reviewing it is itself the harm: an exposed child, a scam that already collected money, a stream nobody removed before it was screenshotted and shared elsewhere.
Keyword lists and simple image hashing catch known, unchanged content but miss context, new phrasing and anything in live video or audio. Models that classify text, images, audio and video against a platform's own policy change what is possible: a violation can be scored and actioned within a second of appearing, at a volume no human team could review, while the decisions that affect someone's account permanently still go through a person.
How does it work?
- Ingest content as it is created. Frames from live video, chat messages, images and audio are pulled into the moderation pipeline as they are posted or streamed, not on a delay.
- Classify against the policy. Vision, audio and language models score each item against the platform's own violation categories (nudity, violence, harassment, CSAM, grooming, self harm, scams) and return a confidence per category.
- Act by severity and confidence. The highest confidence, highest severity matches are actioned automatically (blur, mute, end a stream, remove a message); everything else is queued for a human, ranked by risk.
- Give reviewers the reasoning. A human moderator sees the flagged content, the category and the confidence score, so they decide in seconds instead of watching a whole stream again from the start.
- Enforce and log. Confirmed violations trigger the platform's own enforcement (strike, mute, ban) with an auditable record of what was actioned, by what and when.
- Support appeals. A user can contest a decision; a different reviewer than the original check looks at it again, and confirmed false positives feed back into the category thresholds.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Mainstream
- Channels
- API and system to system, Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | at least 1.5 million | 1 | 1 vendor |
Value drivers: Risk and loss reduction, Compliance quality, Customer experience, Lower cost to serve.
Indicative value
A social, dating or gaming platform with 10 million monthly active users
USD 1.5 million to USD 6.4 million
Manual review cost avoided per year
How this is calculated
Formula: contentItems * shareAutomated * reviewCostPerItem. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| User generated content items needing a moderation decision per year contentItems, items per year | 50,000,000 | 50,000,000 | Editorial assumption for a platform of this scale, replace with your own content volume. |
| Fully loaded human review cost per item reviewCostPerItem, USD per item | 0.05 | 0.15 | Editorial assumption for outsourced or in house moderation review, replace with your own. |
| Share of moderation decisions made without a human review shareAutomated, fraction of items | 0.6 | 0.85 | Editorial assumption, replace with your own; ambiguous content, appeals and any account level action still need a person. |
What it leaves out: Counts only the human reviewer time saved per item on decisions the AI can make alone. It leaves out the cost of the moderation platform itself, the time spent on appeals and escalations, and the larger value: legal exposure, regulatory fines and user trust protected by removing harmful content faster than a human only team could.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Plato
Global · Media and entertainment · 2021
Plato, a social gaming community where members play more than 45 games with one another, adopted Hive as its primary moderation solution as its user base grew. Hive analyses profile pictures and group chats against Plato's community guidelines, with an automatic text filter deployed in public group chats. Hive's own case study reports that Plato reduced user complaints of inappropriate content by more than 90%, especially in public group chats after it deployed that automatic filter.
No outcome disclosed.
Tango
Global · Media and entertainment · 2021
Tango, a live streaming platform with more than 500 million registered users, modernized its legacy in house moderation system with Hive's Vision Language Model and Visual and Text Moderation models to analyse live video streams and chat in real time. The vision language model reads multiple frames of a stream together to catch violations that single frame detection misses, and a text model covers harassment and solicitation in live chat, through a single API. Hive's own case study reports that, since partnering with Hive, Tango has reduced users' exposure to harmful content by more than 80%.
No outcome disclosed.
Chatroulette
Global · Media and entertainment · 2020
Chatroulette, the random video chat service launched in 2009, partnered with Hive in early 2020 as part of a major platform relaunch to implement automated moderation of its live video streams. Hive's visual moderation model flags and closes unsafe streams, and the company's Chief Executive Officer describes user complaints about inappropriate content falling from hundreds a day to fewer than one a week after integration. Hive's own case study reports that the incidence rate of inappropriate content on Chatroulette fell by over 95% in the three months after integration.
- Interactions handled: at least 1.5 million, unsafe streams closed per month
"Hive has helped Chatroulette close over 1.5 million unsafe streams a month with our best-in-class visual moderation model."
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A written, versioned content policy with labelled examples per violation category and severity
- A dataset of past moderation decisions, including confirmed false positives and appeal outcomes
- Age and region flags where the policy differs by jurisdiction or user age
Systems to integrate
- Content and chat pipelines that carry video, image, text and audio as it is created
- A moderation review queue for human moderators, ranked by risk
- The platform's own enforcement systems (mute, strike, temporary or permanent restriction)
- Analytics and regulator transparency reporting
Complexity: Medium
Classifying known violation types in text and static images is mature. The effort is in real time video and audio at low latency, tuning thresholds so obvious violations are actioned automatically without over removing borderline but legitimate content, and building the human review and appeals path before launch, not after an incident forces one.
- 1
Start from your written policy, not a generic model
Translate your own content policy into labelled examples per category and severity. A model tuned on someone else's definition of a violation will both over and under flag against yours.
- 2
Set automation thresholds by severity, not by content type
Auto action only the highest confidence, highest severity matches; route medium confidence and borderline cases to a ranked human queue, and leave low risk content alone.
- 3
Build the review queue and the appeal path before launch
Reviewers need the flagged content, the model's stated category and confidence, and one click actions. Users need a way to contest a decision: moderation without an appeal path erodes trust, and for an online platform within the scope of the EU Digital Services Act, does not meet the Article 20 internal complaint handling requirement.
- 4
Instrument for evasion, not only accuracy
Track how violation patterns shift after every model update, since people adapting to evade detection change tactics quickly, and keep a manual override channel for new violation types the model has not seen yet.
- 5
Localise thresholds and escalation per region
What counts as a violation, and who must review it, varies by jurisdiction and the age of the people affected. Do not run one global threshold across markets with different rules.
Guardrails
- Human review before any permanent account level sanction; only reversible actions (mute, hide, temporary restriction) are fully automated
- Confidence and severity thresholds tuned per category so only high confidence matches are auto actioned
- An appeals path reviewed by a different person than the original decision, logged and reportable
- Regular audits of false positive and false negative rates, per category and per language
KPIs to instrument
- Detection rate and false positive rate per violation category and language, on a labelled audit sample
- Time from content posted to a violation being actioned
- Appeal volume and the share of automated decisions overturned on appeal
- Reviewer caseload and time to decision in the human queue
Human in the loop
Human moderators own every irreversible enforcement action, every appeal and any decision involving a minor or a law enforcement referral. They also review a sampled set of automated decisions every week to catch drift, and confirmed false positives adjust the category thresholds, not just the individual case.
Common failure modes
- Confident wrong removals
- The model removes borderline but legitimate content with high confidence, angering users and inviting regulatory scrutiny. Keep an appeals path with a different reviewer, and audit false positives per category on a schedule, not only after a complaint.
- Evasion drift
- People adjust language, images or stream behaviour to slip under thresholds within days of a model update. Retrain and tune thresholds on a schedule, and treat a rising queue of unclear content as a signal, not noise.
- One threshold for every market
- A single global policy misses local legal definitions of harmful content and who counts as a minor, creating regulatory exposure in some regions. Localise thresholds and escalation paths per jurisdiction.
What are the risks and rules?
EU AI Act
Minimal risk
Classifying user generated content against a platform's own policy is not listed in Annex III, and the system does not generate or manipulate content, so it carries no use case specific duty under the EU AI Act. Platform content moderation is instead the direct subject of the EU Digital Services Act (Regulation (EU) 2022/2065), a separate regulation. Article 17 requires a statement of reasons for content moderation decisions, and Article 20 requires an online platform within its scope to offer an internal complaint handling system, though Article 19 exempts micro and small enterprises from that duty; see the cited guidance below.
Rules that apply
Guidance
- Digital Services Act: statement of reasons and internal complaint handling (European Union, Europe). Regulation (EU) 2022/2065. Article 17 requires providers of hosting services to give a statement of reasons for content moderation decisions. Article 20 requires online platforms to offer an internal complaint handling system for those decisions, but Article 19 exempts micro and small enterprises from that duty.
Controls to put in place
- A documented, versioned content policy with an owner and a change log
- False positive and false negative audits per category and language on a fixed schedule
- An internal appeal path for every enforcement action, reviewed by a different person
- Human review before any permanent account level sanction
Frequently asked questions
- How much can AI reduce harmful content on a platform?
- It depends on the starting point and the content type. Hive Moderation reports that Chatroulette reduced the incidence rate of inappropriate content by over 95% in the three months after integration, Plato reduced user complaints of inappropriate content by more than 90%, and Tango reduced users' exposure to harmful content by more than 80% after continuously deploying new models. These are single vendor reported figures for specific platforms; measure your own before and after on your own content mix.
- Should every moderation decision be automated?
- No. Automate the highest confidence, highest severity matches, and route anything ambiguous, or any permanent account action, to a human. An appeal path reviewed by a different person than the original decision is not optional: without one, users have no recourse from a wrong call.
- Is AI content moderation high risk under the EU AI Act?
- Content moderation of user generated content is not listed in Annex III, so it is not high risk under the AI Act itself. It is however the direct subject of the EU Digital Services Act (Regulation (EU) 2022/2065), a separate regulation. Article 17 requires a statement of reasons for moderation decisions, and Article 20 requires an online platform within its scope to offer an internal complaint handling system, though Article 19 exempts micro and small enterprises from that duty.
- What causes false positives in AI content moderation, and how should they be handled?
- A model can flag borderline but legitimate content with high confidence, especially content that resembles a violation out of context. Tune confidence and severity thresholds per category, audit false positive and false negative rates on a labelled sample on a fixed schedule rather than only after a complaint, and give every user an appeal path reviewed by a different person than the original decision.
How to cite this page
Blits.ai AI Use Case Library, "AI content moderation for trust and safety at online platforms", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/trust-and-safety-content-moderation. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 30 September 2026: First published