paper · Proceedings: 1999

Policy invariance under reward transformations: Theory and application to reward shaping

A 1999 reward-shaping result preserves optimal policies under stated conditions, while leaving the original objective to be justified.

Andrew Y. Ng, Daishi Harada and Stuart Russell ask when extra training rewards preserve optimal policies. Reward shaping can ease learning while accidentally changing the task.[1]

A conditional guarantee

For the finite-state Markov decision processes specified in sections 2–3, adding F(s,a,s′) = γΦ(s′) − Φ(s) preserves optimal policies. Here γ is the discount factor and Φ a potential over states. Undiscounted tasks require proper termination; infinite-state extensions need additional regularity. Theorem 1’s converse means non-potential shaping can fail in some environment, not that it fails in every task.[1]

The introduction recounts a simulated bicycle exploiting progress bonuses by circling; its underlying experiment was not independently reviewed.[1]

For Reward Hacking, this is an editorial historical connection: preserving an objective does not establish that it captures human intentions. Reinforcement Learning: A Survey separately examines objective criteria and learning-period performance.

What the experiments show

Section 4 uses stochastic 10×10 and 50×50 navigation grids with a cost of one per step. A distance-to-goal potential and a half-strength variant reduce steps to the goal during learning compared with no shaping (Figure 1). Results average 40 independent runs; the stated learner is Sarsa with 0.10-greedy exploration and learning rate 0.02.[1]

A second task requires visiting ordered subgoals in a 5×5 grid. Its state includes which flags have been collected: progress cannot be determined from location alone. Both a coarse subgoal potential and a more detailed estimate of remaining travel time improve the displayed learning curves, with the latter performing better (Figure 2).[1]

These examples illustrate how domain knowledge can supply a useful potential. The policy-invariance theorem does not guarantee a learning-speed improvement for every potential, learner or environment. The plots measure steps to a specified goal, not human welfare or safety during exploration.[1]

Historical context and review limits

Russell’s author bibliography records ICML 1999; only the year enters the chronology. Selected original sections 1–3 and 5 were previously checked; section 4 and Figures 1–2 were subsequently inspected on rendered author-copy pages 6–8. Full proofs, code, uncertainty analysis and independent replication were not audited. Mathematical glyphs and the Stanford copy’s text extraction are unreliable; visual inspection supports the experiment description.[1]

Explore the chronology →

Sources

  1. Policy invariance under reward transformations: Theory and application to reward shaping · Source record src-177 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7

Pages that link here

Last updated 2026-10-10