paper · Proceedings: 2000
Algorithms for Inverse Reinforcement Learning
Ng and Russell infer reward functions that make observed decisions optimal, while exposing ambiguity in that inference.
Andrew Y. Ng and Stuart Russell study how to infer a reward function from an agent’s observed decisions. In their Markov decision process formulation, the reward should make the observed policy optimal. The problem reverses ordinary reinforcement learning, which starts with a reward and seeks a policy (sections 1–3).[1]
Approach and ambiguity
The paper characterizes the compatible reward functions for a known policy in a finite, known environment. Behavior alone leaves many solutions: even an everywhere-zero reward makes every policy optimal. The authors therefore add selection criteria, favoring rewards that separate demonstrated actions from alternatives, with an optional penalty favoring simpler rewards. These criteria choose among explanations; they do not establish a uniquely recovered human objective (section 3).[1]
Three algorithms cover tabulated finite-state rewards, linear combinations of fixed features in large state spaces, and finite observed trajectories. Demonstrations use small navigation and mountain-car problems. The authors leave noisy measurements, suboptimal behavior, experimental identifiability and partial observability as open questions (sections 1 and 7).[1]
Alignment relevance and limits
For Outer Alignment, learning an objective from demonstrations offers an alternative to writing every tradeoff manually. This is an editorial connection to objective specification: the paper is general learning and control research, without a demonstration that inferred rewards capture defensible human values.
Apprenticeship Learning via Inverse Reinforcement Learning later cites this work and asks whether good task performance can be obtained despite uncertainty about the true reward.
Historical context
Published in the proceedings of the Seventeenth International Conference on Machine Learning in 2000, as listed in the author’s publication archive. Only the year is used in the chronology. The author-hosted PDF was checked visually on pages 1–4 and 7–8; its text extraction is garbled. Full algorithm derivations, proofs and experiment replication were not audited.[1]
Later public explanation
In his 2016 Science Friday interview, Stuart Russell illustrated objective inference with a person’s morning coffee-making routine. This is an intuitive teaching example (answer to Flatow’s morality question), not a result demonstrated in this paper.[2]
Sources
- Algorithms for Inverse Reinforcement Learning · Source record src-172 · Back to claim ↑1 ↑2 ↑3 ↑4
- Making the Most of A.I.’s Potential · Source record src-194 · Back to claim ↑1
Pages that link here
Last updated 2026-10-10