concept
Inner Alignment
Matching a learned optimizer's objective to the objective used to train it.
Inner alignment asks whether a learned optimizer’s objective matches its training objective. The question arises when training produces a system that itself searches for high-scoring actions or plans: a mesa-optimizer. It differs from Outer Alignment, which asks whether the training objective captures the intended human goal.[1]
A task learned for the wrong reason
The paper offers a hypothetical maze in which every training exit is red. Rewarding arrival at an exit cannot, in those mazes, distinguish seeking exits from seeking red objects. With blue exits and unrelated red objects, those criteria separate. This illustrates possible generalization failure, not a reported experiment or proof of internal search (section 1.3).[1]
Agreement on training data
The 2019 framework distinguishes robust objective agreement from pseudo-alignment: a learned optimizer can perform well during training because a proxy correlates with the target, its represented target is only approximate, or limitations prevent it from exploiting a mismatch. More capable optimization need not repair those causes of apparent agreement.[1]
Pseudo-alignment need not involve deception. Deceptive Alignment is a further proposed case in which training cooperation serves another objective. The framework does not establish how frequently these mechanisms occur in deployed models.[1]
What evidence would settle the question?
For a particular system, ask separately what objective selected it during training, whether it implements search, and what criterion that search uses. An apparent behavioral goal alone does not establish the internal mechanism. This is a reading checklist, not a diagnostic supplied by the paper.
Read the paper writeup for the three forms of pseudo-alignment and the additional assumptions behind deception. The reading guide places this question after proxy failures so the two are easier to distinguish.
Version and review scope
Selected original arXiv v3 HTML passages in sections 1.1–1.3 and 3.1 were checked for this explanation. The paper writeup retains its earlier selected-PDF review and version history. The maze is a thought experiment; no new model experiment, revision comparison or audit of cited studies was performed. NotebookLM indexed-passage retrieval was unavailable for this round.
Sources
Pages that link here
- AI Alignment concept
- Deceptive Alignment concept
- Distributional Shift concept
- Evan Hubinger person
- Key Papers: A Reading Guide note
- Mesa-Optimization concept
- Outer Alignment concept
- Risks from Learned Optimization in Advanced Machine Learning Systems paper
- Surfing the AI Waves paper
Last updated 2026-10-11