First-pass triage
Classifiers score incoming text, images, and video against your policy so obvious breaches are queued and borderline cases are routed to a human, instead of every item waiting for manual review.
Automated moderation can protect users at a scale no human team can match, but only when the risk, fairness, and regulatory questions are answered first. The Institute of AI is an independent professional body with no moderation software to sell, so the advice serves your users and your obligations rather than a licence fee.
Under the Online Safety Act, keeping illegal content and content harmful to children off a UK service is no longer a matter of goodwill. It is a legal duty enforced by Ofcom, with real penalties for services that fail to act. Under the Act a service’s first legal step is an illegal content risk assessment, with a children’s access assessment where children can reach the service, and every later moderation decision is measured against them. The volume of content on any live platform makes purely manual review impossible, so automated moderation has become part of the answer rather than an optional extra.
At the same time, automation carries its own risks. A model that removes the wrong post silences a legitimate voice; a model that misses the right one leaves a user exposed. Both failures are now visible to regulators, to the press, and to the people affected. The task facing organisation leaders is not whether to use AI here, but how to use it in a way they can stand behind.
That is a governance question as much as a technical one. It calls for clear policy, tested models, honest measurement of who is affected, and a named person who owns the outcome. The Institute of AI exists to help you get that right, without a moderation product to push.
or 10% of qualifying worldwide revenue, whichever is greater, the maximum Online Safety Act penalty Ofcom can levy
pillars an organisation is assessed against: Strategy, Governance, Skills, Implementation, Impact
public Charter, free for any organisation to read and sign
The strongest use cases pair machine speed with human judgement and an audit trail a regulator or an appeal can follow.
Classifiers score incoming text, images, and video against your policy so obvious breaches are queued and borderline cases are routed to a human, instead of every item waiting for manual review.
Risk scoring pushes the most severe or most viral content to the front of the queue and sends the rest to the right specialist team, so reviewer time lands where harm is greatest.
Multilingual models extend moderation to languages and dialects a small team cannot staff around the clock, with native-speaker review kept in the loop for nuance and slang.
Perceptual hashing and vision models detect known illegal material and flag novel graphic content before it spreads, sparing reviewers repeated exposure to the worst of it.
Anomaly and pattern detection surface coordinated abuse, scam waves, and grooming signals that no single reviewer would connect across thousands of separate reports.
The same models sample past decisions to check consistency, catch reviewer drift, and give appeals teams a faster, evidenced starting point for a second look.
Six questions we settle before an automated moderation system touches live user content, and keep answering after it does.
A moderation model that looks accurate in aggregate can still be wrong in the ways that matter most, on the rarest and most serious harms. Precision and recall are measured per harm category before launch, on data that reflects real traffic rather than a clean benchmark, and revalidated whenever the model or policy changes.
of harm categories validated separately before go-live
Models trained on historic decisions can over-flag some dialects, communities, or identity terms and under-protect others. Error rates are measured across groups, reclaimed and in-community language is handled with care, and disparities are treated as defects to fix, not noise to accept.
tolerance for unexplained gaps in error rates between groups
The Online Safety Act’s complaints and terms-of-service duties, and basic fairness, mean a user whose content is removed can be told which rule was broken and how to challenge it. Where a model drives that outcome, the specific policy and the evidence behind the decision travel with it into the notice and the appeal.
plain-language policy reason attached to every enforcement action
Moderation processes personal data at scale, often special-category and sometimes about children. A Data Protection Impact Assessment settles the lawful basis, retention, and access before launch, and where an automated action significantly affects a user, such as terminating an account or cutting off monetisation, a genuine route to human review is provided in line with UK GDPR.
completed before any model touches live user content
Accountability for enforcement cannot be handed to a classifier or an outsourced queue. A named senior owner holds the policy, the thresholds, and the outcomes, and the highest-severity decisions keep a trained human in the loop with the support and welfare that difficult work demands.
accountable senior owner for policy and thresholds
Abuse adapts, coded language shifts, and a model that was accurate at launch drifts as traffic changes. Live monitoring of false positives, false negatives, appeal-overturn rates, and reviewer feedback keeps performance honest and triggers retraining before quality quietly erodes.
drift and appeal-overturn monitoring, not a one-off sign-off
Advice, engineering, and professional standards from a single independent body, with no moderation software to sell.
We help you decide where automated moderation genuinely earns its place, set thresholds you can defend to Ofcom and to your users, and map a governance model to the Online Safety Act and UK GDPR. Independent, so the advice serves your risk, not a vendor sale.
Explore AI consultancyThe Institute of AI engineering practice designs triage, routing, and quality-assurance systems around your policy and your existing controls, builds in the audit trail and human-review routes from the start, and hands the system to your own team with the knowledge to run it.
See AI solutionsAs the UK’s professional body for AI, IoAI accredits AI professionals against its published competency framework, at four levels and with annual CPD, and organisation accreditation plus the UK AI Readiness Charter give your board and your users an independent signal that the systems are governed properly.
Organisation accreditationA short call with the Institute of AI. Plain answers on where automated moderation fits, how to govern it, and what good looks like. No moderation software to sell.
See how the Institute of AI tailors independent, standards-led support to the sector you operate in, from public services to online platforms.
Browse sectorsVendor-free advisory for boards setting an AI position that stands up to regulatory scrutiny and their own risk committees.
Learn moreFive pledges that get your organisation, and Britain, ready for AI. Free to read and free to sign.
Read the Charter