paper · Catalog publication: May 1, 1996
Reinforcement Learning: A Survey
A 1996 survey distinguishes the objective an agent optimizes from the performance and penalties of learning.
- First submission: September 1995
Leslie Pack Kaelbling, Michael L. Littman and Andrew W. Moore survey learning through interaction with an environment. Section 1 separates the reward criterion defining a successful policy from the criterion used to assess learning itself.[1]
What counts as success?
Finite-horizon, discounted and average-reward objectives can favor different actions in the same environment (section 1.2, Figure 2). Section 1.3 distinguishes eventual optimality from speed, performance after a fixed time and regret: reward lost while learning compared with acting optimally from the beginning. Faster convergence can incur larger penalties along the way.[1]
For Outer Alignment, this offers a historical comparison: specifying reward amounts leaves time horizon and treatment of learning-period outcomes unresolved. That connection is editorial; the survey does not establish a modern alignment method. Policy invariance under reward transformations: Theory and application to reward shaping examines a different question: changing training rewards while preserving a given objective.
Historical context and review limits
JAIR records May 1, 1996 publication; the PDF records September 1995 submission. Selected section 1 passages and Figure 2 were reviewed, not the entire survey or its cited experiments.[1]
Sources
- Reinforcement Learning: A Survey · Source record src-176 · Back to claim ↑1 ↑2 ↑3
Pages that link here
Last updated 2026-10-09