paper · Catalog publication: May 1, 1996

Reinforcement Learning: A Survey

A 1996 survey distinguishes the objective an agent optimizes from the performance and penalties of learning.

  • First submission: September 1995

Leslie Pack Kaelbling, Michael L. Littman and Andrew W. Moore survey learning through interaction with an environment. Section 1 separates the reward criterion defining a successful policy from the criterion used to assess learning itself.[1]

What counts as success?

Finite-horizon, discounted and average-reward objectives can favor different actions in the same environment (section 1.2, Figure 2). Section 1.3 distinguishes eventual optimality from speed, performance after a fixed time and regret: reward lost while learning compared with acting optimally from the beginning. Faster convergence can incur larger penalties along the way.[1]

For Outer Alignment, this offers a historical comparison: specifying reward amounts leaves time horizon and treatment of learning-period outcomes unresolved. That connection is editorial; the survey does not establish a modern alignment method. Policy invariance under reward transformations: Theory and application to reward shaping examines a different question: changing training rewards while preserving a given objective.

Historical context and review limits

JAIR records May 1, 1996 publication; the PDF records September 1995 submission. Selected section 1 passages and Figure 2 were reviewed, not the entire survey or its cited experiments.[1]

Explore the chronology →

Sources

  1. Reinforcement Learning: A Survey · Source record src-176 · Back to claim ↑1 ↑2 ↑3

Last updated 2026-10-09