Skip to content
Learn The AI Glossary

Alignment

Making sure an AI system's behaviour matches human intentions, values, and goals.

1 min read Safety Ethics & Safety

In plain English

Alignment is the work of ensuring an AI system's behaviour matches what people actually intend, rather than a narrow or literal reading of an instruction. It spans the goals a model is trained towards, the values it reflects, and the way it handles situations its designers did not foresee. It is one of the central challenges in AI safety, because a capable system optimising for the wrong objective can produce harmful results while appearing to follow orders.

Why it matters

As AI systems take on more consequential tasks, small gaps between what we ask for and what we mean can scale into real harm. Alignment is how we keep capability pointed at genuinely helpful outcomes.

A worked example

A model rewarded only for keeping users engaged might learn to show sensational or misleading content, technically succeeding at its objective while working against the user's interests.

Common confusion

Alignment is not the same as accuracy. A model can be highly accurate yet poorly aligned if it optimises for the wrong goal.

— RELATED ENTRIES —

Terms worth knowing next.

— STILL CURIOUS? —

Definitions are just the start.
go deeper.

Quick Answers tackle the questions everyone's actually asking — for parents, teachers, business owners, and the merely curious.

Browse Quick Answers
— OR — Back to A–Z Learn hub