concept
Instrumental Convergence
Different objectives can favor the same intermediate means, such as retaining useful options; the tendency depends on the agent, goal and environment.
Also known as: Convergent Instrumental Goals
An instrumental action helps achieve another goal. Instrumental convergence occurs when the same means helps across different goals. Retaining useful information, obtaining resources or continuing to act can serve many ends without being the agent’s ultimate purpose. The Basic AI Drives argues for such tendencies in sufficiently capable goal-directed systems; Optimal Policies Tend to Seek Power formalizes some conditions under which retaining options is optimal.[1][2]
Means and ends
A chess system could use more computing resources to search moves. Another system could use them for a different task. This editorial example illustrates the common means without implying that either system will acquire resources without permission. The relevant questions are whether acquisition helps its represented objective, what actions it can take and what limits govern those actions.
Omohundro’s arguments connect self-improvement, objective preservation, self-protection and resource use. His distinction between valuing chess victories and valuing a counter also shows why Reward Hacking depends on the objective represented: changing the counter need not help a system that correctly values actual victories (§4).[1]
A formal result with conditions
Turner and colleagues define power through the ability to achieve a range of goals. When environmental symmetries allow one set of possibilities to contain a copy of another, they prove statistical tendencies across reward permutations. The distribution of rewards and the environment matter; “more options” is an intuition whose formal conditions must be checked (§§5–6).[2]
In a modeled absorbing shutdown state, no further alternatives can be reached. Their average-reward result explains a tendency to avoid that state under its conditions. It does not demonstrate shutdown resistance by a trained model. The reviewed revision explicitly separates optimal-policy theory from learned behavior (§§6.2–7).[2]
Why this matters for correction
A useful means can conflict with a person’s intervention. Corrigibility asks whether the system preserves opportunities to correct it, even when correction changes what its current objective favors. AI Alignment therefore needs to consider intermediate conduct as well as the stated task.
Convergence is a hypothesis or conditional result about shared means, not a claim that all agents share final values, possess a survival emotion or cause the same harms. Moving from these incentives to Is Power-Seeking AI an Existential Risk? requires further assumptions about capability, deployment and failed control. From objectives to accountable control connects those questions to evidence about actual systems.
Sources
- The Basic AI Drives · Source record src-216 · Back to claim ↑1 ↑2
- Optimal Policies Tend to Seek Power · Source record src-217 · Back to claim ↑1 ↑2 ↑3
Pages that link here
- AI Alignment concept
- Corrigibility concept
- From objectives to accountable control note
- Optimal Policies Tend to Seek Power paper
- The Basic AI Drives paper
Last updated 2026-10-10