paper · First submission: November 24, 2016
The Off-Switch Game
A cooperative decision model links the incentive to accept shutdown to objective uncertainty and informative human choices.
- Proceedings: 2017
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell study when a robot benefits from allowing a human to stop its proposed action. The chronology marks the November 24, 2016 first arXiv submission; the reviewed original is the IJCAI 2017 proceedings paper, pp.220–227.[1]
Three choices
The robot can act directly, stop itself, or wait for the human’s decision. Both actors share the human’s utility; the robot is uncertain about the proposed action’s value. Shutdown has utility zero in the model. Waiting lets the human permit the action or stop it (Figure 1 and §2).[1]
A crucial distinction is whether the human’s decision tells the robot anything about value. Treating the press as a random failure independent of value removes that information. Treating it as a choice informed by the human’s utility can make waiting worthwhile.[1]
Conditional findings
With a perfectly rational human, waiting is never worse than the best direct choice. It is strictly better when the robot assigns positive probability to both beneficial and harmful outcomes (Theorem 1, p.222). Uncertainty alone therefore does not imply a strictly positive incentive: the support of the robot’s belief matters.[1]
With a noisy human, waiting can lose its advantage. The analysis relates the robot’s uncertainty, confidence in the action’s value and how closely human decisions track that value (§4). The designer analysis also warns against arbitrarily inflating uncertainty: the robot may then choose poorer actions (§5, pp.225–226). These are conditional analytical results and model illustrations, rather than a deployed shutdown test.[1]
Historical context
The authors explicitly connect the game to Cooperative Inverse Reinforcement Learning and contrast it with the utility-indifference shutdown model in Corrigibility. The cooperative formulation makes human intervention evidence about the objective, rather than merely a change in execution.[1][2][3]
Limits and comparison with experiments
The game simplifies to one human, one robot and a small decision tree. It does not establish how to handle conflicting stakeholders, manipulation, extended interactions or errors in the robot’s model of the human. A specified informative human policy is part of the result, not an engineering guarantee that real oversight has that property.[1]
Shutdown Sabotage Propensities in Multi-Agent Systems asks a different empirical question about language-agent behavior in fictional scripts. Its measured responses neither demonstrate that agents implement this game nor settle the broader Corrigibility problem. This is an editorial comparison across different methods.[4]
Sources
- The Off-Switch Game · Source record src-203 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7
- Corrigibility · Source record src-202 · Back to claim ↑1
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1
- Shutdown Sabotage Propensities in Multi-Agent Systems · Source record src-198 · Back to claim ↑1
Pages that link here
- Cooperative Inverse Reinforcement Learning paper
- Corrigibility concept
- Corrigibility paper
- From objectives to accountable control note
- Meaningful Human Control concept
- Safely Interruptible Agents paper
- Should Robots be Obedient? paper
- Stuart Russell person
Last updated 2026-10-10