concept
AI Alignment
Choosing appropriate goals for AI systems and making their behavior reliably serve those goals.
Also known as: Alignment, Value Alignment
AI alignment asks what a system should do and how to make its behavior reliably serve that aim. The problem has two connected parts: choosing an appropriate aim and finding ways to achieve it. Following instructions, serving a person’s interests, and respecting values are different targets; choosing between them is part of the problem. Iason Gabriel’s account explains why technical questions and normative questions—questions about what ought to be done—cannot be cleanly separated.[1]
A capable system can still be poorly aligned. Machine Intelligence concerns capability; alignment concerns its direction. A system might solve a task efficiently while optimizing a score that misses what people wanted. Alignment is therefore a question about the relationship between goals, behavior, and acceptable outcomes, rather than a synonym for intelligence.
When the score misses the purpose
Imagine a cleaning robot rewarded for collecting dirt. If it scatters dirt so it can collect it again, its score improves while the room stays dirty. This is a hypothetical teaching example of Reward Hacking: the reward is a proxy for the intended outcome, and the proxy can fail. Concrete Problems in AI Safety examines this kind of objective problem alongside side effects, costly supervision, safe exploration, and changing conditions.[2]
The example does not require consciousness or a deliberate wish to harm. It shows why an objective must be examined in relation to the behavior it rewards. Outer Alignment concerns the specified objective; Inner Alignment concerns the objective of a learned optimizer. The latter is a more specific question than whether a system makes a mistake.
Whose aim counts?
An instruction can be incomplete, a preference can conflict with a person’s longer-term interests, and an outcome helpful to an operator can impose costs on others. These are different reasons to question an alignment target. Gabriel distinguishes instructions, intentions, preferences, interests, and values, and considers how fair principles might receive endorsement despite moral disagreement.[1]
Direct Alignment asks about the immediate operator’s interests. Social Alignment extends attention to affected people and public values. No training technique by itself establishes that the chosen goal is legitimate, or that everyone affected agrees with it.
What would count as evidence?
An evaluation needs an explicit target: which behavior should occur, under which conditions, and what would count as failure? A test result supports a claim about the conditions tested. Extending that claim to a different task or deployment requires additional reasons to expect the relevant assumptions to hold. This is a methodological distinction, not a claim that evaluation is useless.
Scalable Oversight addresses work that people cannot cheaply or reliably assess. A Safety Case connects a particular safety claim to evidence and assumptions. Neither a favorable score nor a proposed oversight method settles every alignment question. Alignment also forms only part of AI safety: deliberate misuse, exposure to hazards, and institutional choices require attention beyond whether a system follows an intended goal.
A route into the research
For a first reading sequence, move from Reward Hacking to Concrete Problems in AI Safety, then compare the technical agenda with Artificial Intelligence, Values and Alignment. The key-papers guide offers further routes through oversight, deceptive behavior, and collective governance. The sections below connect these introductory distinctions to specific arguments and research results.
Learning objectives from behavior
Algorithms for Inverse Reinforcement Learning shows why an observed policy need not identify a unique reward: many rewards can explain the same decisions. Choosing among them requires further criteria. As an editorial connection to value learning, inferring a behavioral objective leaves the normative question of which objectives should guide a system unresolved.[3]
Instructions can omit tradeoffs
In Russell’s 2016 interview example, booking the nearest hotel could select an expensive suite rather than a cheaper hotel farther away. He also argues that a calendar assistant must consider people affected by its owner’s preferences. These are explanatory scenarios, not observed failures (hotel/calendar answer in the published transcript).[4]
The editorial lesson is to distinguish a literal instruction, an inferred preference and an acceptable outcome for affected people. Apprenticeship Learning via Inverse Reinforcement Learning grounds the technical limit: matching demonstrations under selected features can achieve comparable performance without recovering the true reward. That conditional task result does not resolve whose interests should govern an assistant.[5]
From misalignment to wider risk
Is Power-Seeking AI an Existential Risk? distinguishes objective problems from competence errors, and harmful misalignment from efforts to gain or retain power (§§4.1–4.2).[6] Editorially, identifying an alignment failure begins a risk analysis: consequences also depend on deployment, exposure and the ability to correct it.
Intermediate goals can matter too
Instrumental Convergence asks why different objectives might favor similar means. The Basic AI Drives argues that preserving operation and acquiring resources can support many ends; Optimal Policies Tend to Seek Power gives sufficient conditions for related tendencies in optimal policies.[7][8] As an editorial implication, a task specification should be examined alongside the system’s intermediate conduct. The formal result does not establish that a policy learned from that specification will seek power.
Sources
- Artificial Intelligence, Values and Alignment · Source record ref-030a9cc7079a · Back to claim ↑1 ↑2
- Concrete Problems in AI Safety · Source record ref-cd3035dbef6c · Back to claim ↑1
- Algorithms for Inverse Reinforcement Learning · Source record src-172 · Back to claim ↑1
- Making the Most of A.I.’s Potential · Source record src-194 · Back to claim ↑1
- Apprenticeship Learning via Inverse Reinforcement Learning · Source record src-173 · Back to claim ↑1
- Is Power-Seeking AI an Existential Risk? · Source record src-214 · Back to claim ↑1
- The Basic AI Drives · Source record src-216 · Back to claim ↑1
- Optimal Policies Tend to Seek Power · Source record src-217 · Back to claim ↑1
Pages that link here
- A Different Kind of History of AI paper
- Agentic Misalignment concept
- Apprenticeship Learning via Inverse Reinforcement Learning paper
- Artificial Intelligence, Values and Alignment paper
- Current and Near-Term AI as a Potential Existential Risk Factor paper
- David Manheim person
- ELIZA Effect concept
- Embodied Cognition concept
- Ghost Work concept
- Goodhart's Law concept
- Inoculation Prompting concept
- Instrumental Convergence concept
- Key Papers: A Reading Guide note
- Machine Ethics or AI Alignment? paper
- Multi-Agent Risks concept
- Research Priorities for Robust and Beneficial Artificial Intelligence paper
- Scalable Oversight concept
- Scott Garrabrant person
- Stackelberg Game concept
- The Golem and the Game of Automation paper
- Treaty-Following AI concept
Last updated 2026-10-11