paper · Proceedings: 2004

Apprenticeship Learning via Inverse Reinforcement Learning

Abbeel and Ng match demonstrated feature expectations to obtain comparable task performance without identifying the true reward.

Pieter Abbeel and Andrew Y. Ng address tasks whose reward tradeoffs are difficult to specify, using expert demonstrations instead. Building on Algorithms for Inverse Reinforcement Learning, they shift attention from identifying the reward to performing well under the expert’s unknown reward (sections 1–3).[1]

Approach and guarantee

The reward is assumed to be a weighted sum of known, bounded features. The learner estimates the expert’s expected discounted feature totals and alternates between choosing candidate reward weights and solving the corresponding reinforcement-learning problem. Matching these feature expectations bounds the difference in expected return for normalized weights, even if the true weights are never recovered. The expert need not itself be optimal (sections 2–4).[1]

The guarantee depends on the reward representation, demonstration estimates and access to an appropriate reinforcement-learning solver. The paper describes selecting a policy with human inspection or constructing a mixture of learned policies; it does not guarantee that every intermediate policy performs well. Finite demonstration samples introduce estimation error.[1]

Evidence and limits

Experiments use gridworlds and simulated highway driving. Driving demonstrations include deliberately unsafe styles, which the algorithm also attempts to imitate. The driving experiment has no specified ground-truth reward, so the authors report feature expectations and qualitative imitation rather than measured return under that unknown reward (section 5.2).[1]

For AI Alignment, the editorial lesson is that matching a demonstrator’s behavior under selected features is distinct from deciding whether its objectives are desirable. The simulation evaluates imitation, not real-world driving safety or moral legitimacy.

Historical context

Published at the Twenty-first International Conference on Machine Learning in 2004; the original PDF and author archive identify that venue and year. The chronology uses year precision. Sections 1–6 and the driving table were inspected; appendix proofs, supplementary material, code and replication were not audited.[1]

Explore the chronology →

Sources

  1. Apprenticeship Learning via Inverse Reinforcement Learning · Source record src-173 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5

Last updated 2026-10-09