Skip to content
Learn The AI Glossary

Red-teaming

Deliberately testing an AI system by trying to make it fail or misbehave.

1 min read Safety Ethics & Safety

In plain English

Red-teaming is the practice of deliberately testing an AI system by trying to make it fail, produce harmful outputs, or behave unexpectedly, in order to find vulnerabilities before deployment. It borrows the idea from security, where a friendly team plays the attacker. The findings then feed back into guardrails and design.

Why it matters

Systems behave differently under adversarial pressure than in friendly testing. Red-teaming surfaces the failure modes that real users, and bad actors, will eventually find.

A worked example

A team tries dozens of phrasings to trick a customer-service bot into giving banned advice, then uses what works to strengthen its guardrails.

Common confusion

Red-teaming is not ordinary quality testing. It actively tries to break the system rather than confirming it works under normal use.

— RELATED ENTRIES —

Terms worth knowing next.

— STILL CURIOUS? —

Definitions are just the start.
go deeper.

Quick Answers tackle the questions everyone's actually asking — for parents, teachers, business owners, and the merely curious.

Browse Quick Answers
— OR — Back to A–Z Learn hub