paper · First submission: June 9, 2016
Cooperative Inverse Reinforcement Learning
A shared-reward game makes teaching and asking for information part of learning to assist a human.
- Proceedings: 2016
- Revision: February 17, 2024
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell formulate value alignment as a cooperative game: the human knows a reward parameter that the robot initially does not, while both receive the human’s reward. The chronology marks June 9, 2016 first arXiv submission. This entry reviews the NIPS 2016 proceedings original; the arXiv record also contains a 2024 revision, whose changes have not been compared.[4]
Why interaction changes the problem
In ordinary inverse reinforcement learning, a learner interprets behavior as evidence of a reward. CIRL also makes the demonstrator’s choices part of the joint decision problem. A human can sacrifice immediate task reward to teach something that improves the robot’s later decisions. The robot assists the human’s interests rather than merely adopting the human’s personal desires (§1).[4]
The paper reduces computation of an optimal joint policy to a partially observable Markov decision process, and identifies a sufficient statistic based on the robot’s belief about the reward parameter (§3). This is a formal reduction, not evidence that arbitrary real-world assistance is easy to compute.[4]
Teaching versus performing
An office-supply example gives the human and robot different production capacities. A demonstration chosen only for the human’s immediate reward can communicate less useful preference information than one selected for the joint outcome (§3.3).[4]
The experiments use a discrete navigation grid with linear rewards. They compare expert task demonstrations with an approximate best-response teaching policy at three and ten reward features, testing each condition on 500 sampled reward parameters. The authors report lower regret for the teaching policy: the robot’s later policy loses less reward relative to knowing the true parameter (§4, pp.7–8). Figure 1 illustrates a teacher visiting two valuable regions while an expert stays at the best one. These are computational experiments with modeled demonstrators, not a human-participant or deployed-robot study.[4]
Historical context
The paper explicitly builds on inverse and apprenticeship learning, including Algorithms for Inverse Reinforcement Learning and Apprenticeship Learning via Inverse Reinforcement Learning. The Off-Switch Game subsequently uses this cooperative formulation to study the incentive to permit human intervention. The connection to Outer Alignment is editorial: the target becomes assistance under a specified human reward model, while choosing that model remains consequential.[4]
Limits
CIRL assumes a common payoff, a specified reward family and a human who knows the relevant reward parameter. These assumptions do not solve disagreement among people or ensure that the model represents what affected people value. Teaching also depends on the learner’s interpretation of behavior. The paper’s algorithms and navigation results establish claims within their games; they do not demonstrate general value alignment or a complete solution to Corrigibility.[4]
A 2026 tractable special case
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response develops a special case in which teaching resolves uncertainty in one step without sacrificing task value. It also identifies a failure of literal goal inference when an action is optimal for more than one goal. This extends CIRL’s interactive interpretation while retaining explicit restrictions on the goal and action spaces and the human model.[3]
When the teaching assumption is wrong
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning compares matched and mismatched interpretations of human behavior. Its formal ordering permits a correctly matched teaching pair to perform best while making the reverse mismatch especially poor. Reanalysis of human demonstrations exposes further departures from the modeled teacher. CIRL’s benefit from interaction therefore leaves an important practical question: how well does the assistant’s interpretation fit this person and task?[2]
Learning to coordinate over repeated interactions
Provably Optimal Learning Algorithms for Assistance Games gives efficient online algorithms for a finite repeated-game variant. Its benchmark is the best fixed human–assistant policy pair in hindsight, reduced by an approximation factor. The contribution makes coordination costs explicit, while assuming an algorithmic human with full-information feedback and preferences fixed before play. This develops CIRL’s computational question without establishing that an ordinary person will follow the required learning procedure.[1]
Sources
- Provably Optimal Learning Algorithms for Assistance Games · Source record src-209 · Back to claim ↑1
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning · Source record src-207 · Back to claim ↑1
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Source record src-206 · Back to claim ↑1
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7
Pages that link here
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response paper
- From objectives to accountable control note
- Human Model Misspecification concept
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning paper
- Provably Optimal Learning Algorithms for Assistance Games paper
- Stuart Russell person
- The Off-Switch Game paper
Last updated 2026-10-10