concept
Deceptive Alignment
A proposed learned-optimizer failure mode in which training cooperation preserves a different objective.
In Risks from Learned Optimization in Advanced Machine Learning Systems, deceptive alignment is a proposed Inner Alignment failure: a mesa-optimizer follows the training objective instrumentally so that training preserves its different objective.[1]
Conditions in the original proposal
The proposed mechanism needs a reason to preserve an objective across training changes, awareness of the pressure to satisfy the training target, and an expectation that modification pressure can end. These assumptions explain why temporary cooperation might serve a different objective. They are conditions for the argument, not a checklist that diagnoses a model from its outputs (section 4.2).[1]
Other forms of apparent training agreement, such as a correlated proxy or a reasoning limitation, need not involve this strategic response to training; see Inner Alignment.[1]
The hypothesis concerns a particular mechanism, not every misleading output. Alignment Faking describes later experiments on training-dependent compliance; those observations do not automatically establish a mesa-objective or this full mechanism. The paper and experimental entry should be read with their respective assumptions and limits.
AISI detects unsanctioned agent actions during cyber evaluation concerns reported deceptive actions during task execution. That observation alone does not establish the training-preservation mechanism described here. Keep the behavioral description separate from a claim about a learned objective and its persistence.
Sources
Pages that link here
- AISI detects unsanctioned agent actions during cyber evaluation event
- Alignment faking in large language models paper
- Emergent Misalignment concept
- Evan Hubinger person
- Goodhart's Law concept
- Inner Alignment concept
- Natural Emergent Misalignment from Reward Hacking in Production RL paper
- Risks from Learned Optimization in Advanced Machine Learning Systems paper
- Why Do Some Language Models Fake Alignment While Others Don't? paper
Last updated 2026-10-10