Skip to content
— AI use case —

AI for content moderation,
governed honestly.

Automated moderation can protect users at a scale no human team can match, but only when the risk, fairness, and regulatory questions are answered first. The Institute of AI is an independent professional body with no moderation software to sell, so the advice serves your users and your obligations rather than a licence fee.

— State of play —

Moderation is now a regulated duty.

Under the Online Safety Act, keeping illegal content and content harmful to children off a UK service is no longer a matter of goodwill. It is a legal duty enforced by Ofcom, with real penalties for services that fail to act. Under the Act a service’s first legal step is an illegal content risk assessment, with a children’s access assessment where children can reach the service, and every later moderation decision is measured against them. The volume of content on any live platform makes purely manual review impossible, so automated moderation has become part of the answer rather than an optional extra.

At the same time, automation carries its own risks. A model that removes the wrong post silences a legitimate voice; a model that misses the right one leaves a user exposed. Both failures are now visible to regulators, to the press, and to the people affected. The task facing organisation leaders is not whether to use AI here, but how to use it in a way they can stand behind.

That is a governance question as much as a technical one. It calls for clear policy, tested models, honest measurement of who is affected, and a named person who owns the outcome. The Institute of AI exists to help you get that right, without a moderation product to push.

— Signals —
£18m

or 10% of qualifying worldwide revenue, whichever is greater, the maximum Online Safety Act penalty Ofcom can levy

Five

pillars an organisation is assessed against: Strategy, Governance, Skills, Implementation, Impact

One

public Charter, free for any organisation to read and sign

— Where AI helps —

Where AI helps with content moderation.

The strongest use cases pair machine speed with human judgement and an audit trail a regulator or an appeal can follow.

First-pass triage

Classifiers score incoming text, images, and video against your policy so obvious breaches are queued and borderline cases are routed to a human, instead of every item waiting for manual review.

Prioritisation and routing

Risk scoring pushes the most severe or most viral content to the front of the queue and sends the rest to the right specialist team, so reviewer time lands where harm is greatest.

Language and context coverage

Multilingual models extend moderation to languages and dialects a small team cannot staff around the clock, with native-speaker review kept in the loop for nuance and slang.

Image, audio, and video screening

Perceptual hashing and vision models detect known illegal material and flag novel graphic content before it spreads, sparing reviewers repeated exposure to the worst of it.

Emerging-harm detection

Anomaly and pattern detection surface coordinated abuse, scam waves, and grooming signals that no single reviewer would connect across thousands of separate reports.

Appeals and quality assurance

The same models sample past decisions to check consistency, catch reviewer drift, and give appeals teams a faster, evidenced starting point for a second look.

— What must be governed —

Speed is easy. Answerability is the work.

Six questions we settle before an automated moderation system touches live user content, and keep answering after it does.

01
— Model risk and validation —

A classifier is a claim you must test

A moderation model that looks accurate in aggregate can still be wrong in the ways that matter most, on the rarest and most serious harms. Precision and recall are measured per harm category before launch, on data that reflects real traffic rather than a clean benchmark, and revalidated whenever the model or policy changes.

100%

of harm categories validated separately before go-live

02
— Fairness and bias —

Enforcement must not fall unevenly

Models trained on historic decisions can over-flag some dialects, communities, or identity terms and under-protect others. Error rates are measured across groups, reclaimed and in-community language is handled with care, and disparities are treated as defects to fix, not noise to accept.

0

tolerance for unexplained gaps in error rates between groups

03
— Explainability —

Every removal needs a reason a person can read

The Online Safety Act’s complaints and terms-of-service duties, and basic fairness, mean a user whose content is removed can be told which rule was broken and how to challenge it. Where a model drives that outcome, the specific policy and the evidence behind the decision travel with it into the notice and the appeal.

1

plain-language policy reason attached to every enforcement action

04
— Data protection —

UK GDPR applies to the content and the reviewers

Moderation processes personal data at scale, often special-category and sometimes about children. A Data Protection Impact Assessment settles the lawful basis, retention, and access before launch, and where an automated action significantly affects a user, such as terminating an account or cutting off monetisation, a genuine route to human review is provided in line with UK GDPR.

DPIA

completed before any model touches live user content

05
— Human accountability —

The model advises, a named owner decides

Accountability for enforcement cannot be handed to a classifier or an outsourced queue. A named senior owner holds the policy, the thresholds, and the outcomes, and the highest-severity decisions keep a trained human in the loop with the support and welfare that difficult work demands.

1

accountable senior owner for policy and thresholds

06
— Monitoring after go-live —

Harm evolves faster than a model

Abuse adapts, coded language shifts, and a model that was accurate at launch drifts as traffic changes. Live monitoring of false positives, false negatives, appeal-overturn rates, and reviewer feedback keeps performance honest and triggers retraining before quality quietly erodes.

Ongoing

drift and appeal-overturn monitoring, not a one-off sign-off

Moderate at scale,
answer for every call.

A short call with the Institute of AI. Plain answers on where automated moderation fits, how to govern it, and what good looks like. No moderation software to sell.