paper · First submission: June 12, 2017
Deep reinforcement learning from human preferences
A 2017 study learns rewards from comparisons of short behavior clips, scaling human feedback to deep RL while exposing limits of learned proxies and static feedback.
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg and Dario Amodei investigate how to communicate a task when writing its reward function or demonstrating the desired behavior is difficult. They learn a reward model from human comparisons, then use that model to train deep reinforcement-learning agents in Atari games and simulated robotics. Their contribution is scaling an existing family of preference-learning methods, rather than inventing human feedback from scratch.[1]
Learning from comparisons
People compare two clips lasting one to two seconds. They can prefer either clip, mark a tie or decline to compare them; incomparable pairs are excluded from training. A neural reward model predicts preferences using the total predicted reward over each clip. An ensemble supplies both an averaged reward and a rough disagreement signal for selecting queries. Disagreement-based selection is a heuristic and can impair performance on some tasks.[1]
Three processes run asynchronously: the policy generates behavior and learns from predicted rewards, selected clips go to a human, and the reward model learns from the accumulated comparisons. The policy uses advantage actor-critic for Atari and trust-region policy optimization for simulated robotics. People do not label every environment interaction.[1]
This differs from Apprenticeship Learning via Inverse Reinforcement Learning: the supervisory data consists of comparisons of the agent’s behavior, rather than demonstrations from an expert. A person may recognize a desired movement without being able to perform it.[1]
What the experiments show
Benchmark experiments hide the environment’s reward from the learning agent but use it to measure performance. The study compares real human feedback, synthetic preferences from an oracle with access to that reward, and ordinary reinforcement learning trained directly on it. Eight simulated robotics tasks show broadly comparable mean performance with 700 human comparisons, with less stable learned-reward training. Results across seven Atari games vary: learning remains substantially below the direct-reward baseline on some tasks, and human-feedback Qbert fails to beat the first level.[1]
The main figures average five runs for simulated-robotics baselines and synthetic feedback, and three for Atari baselines and synthetic feedback; each human-feedback curve represents a single run. Some robotics feedback comes from authors. Human feedback can also shape the target differently from the environment reward, so a higher score does not isolate more accurate recovery of that reward.[1]
Separate demonstrations train novel behaviors with author feedback, including a simulated Hopper performing repeated backflips using 900 comparisons in under an hour. These are simulated behaviors, not physical-robot deployments. Sparse feedback reduces the labeling burden in these settings; it does not establish that people can judge every important consequence of a harder task.[1]
Feedback must follow changing behavior
An ablation gathers comparisons only at the beginning of training. The resulting reward predictor can miss behaviors later produced by the policy. In Pong, the agent sometimes avoids losing points without trying to score, sustaining repetitive volleys. The authors use this as evidence for continuing feedback as behavior changes.[1]
This connects to Reward Hacking: a learned reward remains a proxy that optimization can exploit. For Scalable Oversight, the editorial implication is to examine both the cost of feedback and the evaluator’s ability to see the relevant outcome. The experiments concern short, recognizable behaviors; they do not validate supervision of systems whose important errors people cannot detect.
Historical context
The first arXiv version was submitted on June 12, 2017, which supplies this article’s single chronology entry. Later revisions are distinct bibliographic stages. InstructGPT subsequently applies preference-based reward learning to broad language tasks, with demonstrations and different policy-training choices. The shared reward-learning pattern should not erase the differences in task, feedback or evaluation.[1]
Evidence scope
The original v1 introduction, methods, main results, ablations and discussion were inspected. Rendered PDF pages 5 and 8 supplied independent checks of the preference model and Atari run-count caption. The full appendices, videos, implementation, cited originals, revision comparison and independent replication were not audited. The original PDF is ready in the research notebook; indexed identity and selected prose were checked without certifying complete extraction fidelity.[1]
Sources
Pages that link here
Last updated 2026-10-10