concept
Corrigibility
Remaining open to correction, including preserving the ability to intervene when an objective or learning rule is flawed.
Corrigibility concerns an agent’s willingness to tolerate or assist correction, including changes to its goals and shutdown. In Soares and colleagues’ formulation, it also includes avoiding manipulation, maintaining safety mechanisms and preserving correction as the system creates successors or changes itself (§1.1, p.2). An obedient response to one command covers only part of this problem.[1]
The authors describe a corrigible agent as one that “tolerates or assists many forms of outside correction” (§1.1, p.2). The concern is an incentive: if changing the objective would reduce success according to the current objective, a capable optimizer can have reason to resist that change. This does not require attributing fear, consciousness or a biological survival drive.[1]
Incentives and actual behavior
Instrumental Convergence explains why preserving the ability to act can serve many goals. Optimal Policies Tend to Seek Power proves a shutdown-avoidance tendency for average-optimal policies under specified environment conditions. Its reviewed revision also stresses that learned policies need not resemble those optimal policies (§§6.2–7).[5] Editorially, a correction claim needs both an account of incentives and evidence that the actual system preserves intervention; neither can substitute for the other.
The shutdown problem as a small model
A button can stop execution when pressed while the agent still has an incentive to obstruct the press. Rewarding shutdown too strongly can create the opposite incentive. The 2015 paper also asks whether agents will preserve shutdown behavior in new subsystems; its examined utility-indifference construction does not make even a small cost of maintaining that behavior worthwhile (§§2–4). These are analyses of simplified utility-maximizing agents, rather than observations of deployed systems.[1]
The Off-Switch Game studies another route: a robot serves a human whose utility it does not fully know, and the human’s decision can inform that uncertainty. A perfectly rational human makes waiting for their decision at least as valuable as acting or stopping directly. With imperfect human decisions, the result depends on the robot’s uncertainty and model of the human (§§3–5).[2]
Interruption and interpreting a correction
Safely Interruptible Agents studies repeated intervention during reinforcement learning. Under specified convergence and exploration assumptions, Q-learning and a modified Sarsa update can learn the ordinary task without developing a lasting intervention-induced bias. This concerns an asymptotic learning property; it assumes the task reward is correct and does not establish a physical guarantee of timely shutdown.[3]
A 2026 assistance-game result examines how a robot interprets human action. An apparently task-optimal action can fit several goals, leaving a literal learner unable to infer the intended one. Modeling the action as a purposeful signal resolves uncertainty in one step for the paper’s special class of games, where every goal has a distinct signal at no task cost. Its notion of responsiveness does not establish openness to changing the learning rule itself.[4]
Correcting the learning rule
Uncertainty is useful only within a specification of what to learn and how to interpret evidence. The 2015 authors explicitly caution that a system can update an uncertain objective while resisting correction of a mistaken rule for determining that objective (§1.1). The off-switch result therefore addresses a specified cooperative game; it does not resolve every form of correction.[1][2]
An editorial comparison with Meaningful Human Control separates two questions: will the agent preserve intervention, and do people have the understanding and responsibility needed to use it well? From objectives to accountable control connects these requirements. Neither a functioning button nor a favorable toy-game result establishes both.
Sources
- Corrigibility · Source record src-202 · Back to claim ↑1 ↑2 ↑3 ↑4
- The Off-Switch Game · Source record src-203 · Back to claim ↑1 ↑2
- Safely Interruptible Agents · Source record src-205 · Back to claim ↑1
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Source record src-206 · Back to claim ↑1
- Optimal Policies Tend to Seek Power · Source record src-217 · Back to claim ↑1
Pages that link here
- Cooperative Inverse Reinforcement Learning paper
- Corrigibility paper
- From objectives to accountable control note
- Human Model Misspecification concept
- Instrumental Convergence concept
- Is Power-Seeking AI an Existential Risk? paper
- Learning under Misspecified Objective Spaces paper
- Meaningful Human Control concept
- Optimal Policies Tend to Seek Power paper
- Provably Optimal Learning Algorithms for Assistance Games paper
- Safely Interruptible Agents paper
- Should Robots be Obedient? paper
- The Basic AI Drives paper
- The Off-Switch Game paper
Last updated 2026-10-10