paper · Proceedings: 2016
Safely Interruptible Agents
A formal reinforcement-learning framework separates temporary human intervention from the task being learned.
- Revision: October 28, 2016
Laurent Orseau and Stuart Armstrong study whether a reinforcement-learning agent can be interrupted repeatedly without learning to avoid or seek the interruption. The paper appeared in UAI 2016. The reviewed MIRI-hosted manuscript explicitly marks an October 28, 2016 revision; the chronology uses the proceedings year because an initial release day has not been established here.[1]
Intervention can change what is learned
In the warehouse illustration, a robot can sort boxes indoors or carry them inside. A human sometimes stops the outdoor activity when it rains. If the learner treats that intervention as part of the ordinary task, it can learn a different preference between activities. The illustration explains an incentive problem; it is not a report of a robot disabling a real switch (Figure 1, p.1).[1]
The framework instead models interruption as temporarily imposing a different policy on the agent. This can mean stopping movement or taking a safer action. The question is whether the agent’s underlying policy learns the task as if future interventions would not occur. The formal property is asymptotic: early learning can be biased even when the bias disappears in the limit (Definition 6, p.4).[1]
Why the update rule matters
Q-learning estimates the value of an action using the best estimated next action, rather than the next action actually executed. Under the paper’s assumptions, this off-policy update can converge to the ordinary task optimum despite interventions. Sarsa normally uses the actual next action, so an imposed intervention changes its update. Safe Sarsa substitutes an action sampled from the underlying policy for that part of the update (§3, Theorems 14–17).[1]
The result requires a finite communicating Markov decision process, bounded rewards, suitable learning rates and continuing exploration of state-action pairs. Interruption depends on the current state in this analysis. The interruption schedule must preserve enough exploration while its reliability approaches one. The paper also gives a theoretical construction for an ideal, uncomputable learner in deterministic computable environments (§4); that construction is not an implementable general safety system.[1]
Historical context
The paper explicitly builds on utility-indifference work discussed in Corrigibility. Its introduction assumes the task reward is correct and asks how interventions avoid corrupting learning. The Off-Switch Game addresses another question: whether a human decision is useful evidence about an uncertain objective. Preventing learning bias and learning from human judgment are different design choices, despite both involving intervention.[1][2][3]
Limits
Safe interruptibility does not by itself establish physical containment, correct objectives, finite-time safety or willingness to accept an arbitrary goal change. The authors leave policy-search algorithms, faster convergence of interruption reliability and scheduled interruptions as open questions (§5). The broader Corrigibility problem also asks whether correction mechanisms survive changes to the agent itself.[1][2]
Sources
- Safely Interruptible Agents · Source record src-205 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7
- Corrigibility · Source record src-202 · Back to claim ↑1 ↑2
- The Off-Switch Game · Source record src-203 · Back to claim ↑1
Pages that link here
Last updated 2026-10-10