paper · First submission: May 28, 2017
Should Robots be Obedient?
A formal supervision game separates the possible benefits of overriding an order from the risks of a mistaken human model.
- Proceedings: 2017
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan and Stuart Russell model a robot learning shared reward parameters from a human’s orders. The chronology marks first arXiv submission; this entry reviews the IJCAI 2017 proceedings original.[1]
Orders as evidence
The human knows the reward but can choose suboptimal actions. Most analysis uses independent rounds, removing exploration and delayed consequences. With the correct model, choosing the action that maximizes expected reward cannot do worse in expectation than obeying; positive advantage requires sometimes departing from the order (§§2–3). This concerns modeled reward, not permission to overrule people.[1]
What the model leaves out
Figure 4 shows that missing reward features can make a simulated learner less obedient and worse than following orders. The paper also distinguishes estimation methods: maximum-likelihood action choices resist misspecifying the noise parameter under its assumptions, while the posterior-mean policy need not. A first-order obedience test provides a restricted missing-feature detector (§§4–5). These are formal results and simulations, not human-participant or deployed safety evidence.[1]
Historical context and limits
The authors connect their analysis to The Off-Switch Game and Safely Interruptible Agents. Human Model Misspecification explains why reward learning alone does not establish Corrigibility: a model can omit precisely the reason for a correction. The paper assumes a shared, static, linear reward; it does not decide authority, rights or disagreement among affected people.[1]
Sources
- Should Robots be Obedient? · Source record src-208 · Back to claim ↑1 ↑2 ↑3 ↑4
Pages that link here
- Human Model Misspecification concept
- Learning under Misspecified Objective Spaces paper
- Outer Alignment concept
- Stuart Russell person
Last updated 2026-10-10