paper · First submission: December 3, 2019
Optimal Policies Tend to Seek Power
Environmental symmetries can make keeping options open optimal across reward permutations; the result does not establish how learned agents behave.
- Proceedings: 2021
- Revision: January 28, 2023
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch and Prasad Tadepalli formalize one part of Instrumental Convergence: when does an optimal policy have reason to retain control over future possibilities? Their finite Markov decision process model makes power relative to a distribution of reward functions, using normalized expected optimal value to capture the ability to achieve different goals (§§3–5).[1]
What makes an option useful to many goals?
If one branch contains a structurally equivalent copy of the possibilities in another, plus further useful possibilities, the authors can compare which actions are optimal as rewards are permuted across states. Proposition 6.9 gives sufficient conditions linking such options to power and action optimality. Simply counting available actions is not the theorem: reachability, the symmetry and additional non-dominated possibilities matter (§6.1).[1]
Here, “most” compares members of each orbit of reward distributions under state permutations (Definition 6.5). It is not a measured proportion of objectives chosen by developers; the authors acknowledge that some permutations may be more likely in practice (p.7, footnote 4; §7).[1]
Why shutdown matters in the model
For average-reward optimal policies, Theorem 6.13 compares sets of long-run state distributions under copy and disjoint-support conditions. Corollary 6.14 shows a tendency to end up outside a particular self-looping state when other such states are reachable. An absorbing shutdown state cannot be left, so avoiding ending there also requires avoiding entry. Figure 8 illustrates this with Pac-Man dynamics while varying the reward, rather than retaining the game’s usual score (§§6.2–6.3, pp.9–10).[1]
This is an incentive result, not an implemented shutdown test. The same conclusion does not forbid passing through an ordinary self-looping state that the agent can leave.
Optimal policies and learned behavior
The reviewed 2023 revision stresses that “learned policies are rarely optimal” (§7, p.10). Reward during training can shape a policy’s generalization without making it a robust reward optimizer. Partial observability and irregular dynamics also limit application; the paper leaves the relationship between power-seeking and disempowering other agents unresolved (§7, pp.10–11).[1]
Editorially, Corrigibility needs evidence about the actual system’s response to intervention. These theorems clarify sufficient conditions for incentives, but do not prove that all capable AI systems seek power or supply a probability of catastrophe. Is Power-Seeking AI an Existential Risk? addresses the additional agent and deployment premises.
Historical context
The first arXiv submission was December 3, 2019; the paper appeared at NeurIPS 2021. This entry uses arXiv v10, revised January 28, 2023, including its explicit learned-policy cautions. Earlier versions and the proceedings text have not been compared in full.[1]
Sources
- Optimal Policies Tend to Seek Power · Source record src-217 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6
Pages that link here
- AI Alignment concept
- Corrigibility concept
- From objectives to accountable control note
- Instrumental Convergence concept
- The Basic AI Drives paper
Last updated 2026-10-10