concept
Mesa-Optimization
Optimization performed by a learned system, whose objective may differ from its training objective.
A mesa-optimizer is a learned system that itself optimizes. Its objective may differ from the training objective, raising Inner Alignment questions. The distinction comes from Risks from Learned Optimization in Advanced Machine Learning Systems. [1]
Internal search and apparent goals
The base objective selects the learned algorithm during training; a mesa-objective guides search performed by that learned algorithm. An objective reconstructed from its behavior need not reveal the criterion implemented internally. Nor does the term require a separate agent inside the model: the learned network itself may implement optimization (section 1.1).[1]
The paper distinguishes internal search from behavior that merely maximizes a score; identifying an apparent behavioral goal is insufficient to diagnose a mesa-optimizer.
Sources
Pages that link here
- Deceptive Alignment concept
- Distributional Shift concept
- Evan Hubinger person
- Inner Alignment concept
- Key Papers: A Reading Guide note
- Risks from Learned Optimization in Advanced Machine Learning Systems paper
Last updated 2026-10-09