paper · First submission: March 4, 2022
Training language models to follow instructions with human feedback
A 2022 study uses demonstrations, human rankings and reinforcement learning to improve GPT-3 instruction following, while exposing limits of labeler preferences and safety.
Long Ouyang and colleagues train GPT-3 models to respond more closely to users’ intended tasks. Their InstructGPT study compares ordinary language-model training with fine-tuning on demonstrations and human preferences across a broad set of language tasks. It treats instruction following as an empirical alignment problem, while explicitly limiting the target to the selected labelers and researchers.[2]
From demonstrations to a learned reward
The procedure has three stages. Contractors first write desired responses, which provide supervised fine-tuning data. They then rank candidate responses to the same prompt; a reward model learns to predict these preferences. Finally, proximal policy optimization (PPO) updates the language model to obtain higher reward-model scores. Most comparison data comes from supervised policies, with some from PPO policies; this study does not continuously obtain a human judgment for every policy update.[2]
A penalty for divergence from the supervised policy is intended to limit reward-model overoptimization. The main InstructGPT variant, PPO-ptx, also mixes in gradients from the original pretraining distribution to reduce losses on other language tasks. These are optimization choices for a learned proxy, not a guarantee that the reward measures all desired outcomes. The distinction connects to Outer Alignment and Reward Hacking.[2]
What the evaluations establish
On the study’s held-out customer prompt distribution, human evaluators prefer the 1.3-billion-parameter InstructGPT model’s outputs to those of the 175-billion-parameter GPT-3 baseline. The paper also reports fewer unsupported additions on closed-domain tasks and improvements on truthfulness evaluations. These results concern response preference and task behavior under the stated evaluation conditions; they do not show that the smaller model is more capable on every task.[2]
Safety results are more qualified. In the RealToxicityPrompts evaluations, InstructGPT produces less toxic text when instructed to respond respectfully; the advantage disappears without that instruction. Explicitly asking for toxic text can make its outputs more toxic than GPT-3’s. The modified Winogender and CrowS-Pairs evaluations do not show reduced bias. Pretraining updates mitigate the regressions caused by PPO fine-tuning, but performance still trails GPT-3 on some tested datasets.[2]
Whose intent supplies the target?
Training uses labeler-written prompts and prompts submitted to early InstructGPT models in the API Playground, not production API customer data in this paper. Splits separate customers by user ID; the tasks are over 96 percent English. About forty contractors provide the demonstrations, comparisons and main evaluations. The authors identify labeler selection, researchers’ instructions and the customer prompt distribution as choices shaping the target behavior.[2]
Training prioritizes helpfulness, whereas final evaluations ask labelers to prioritize truthfulness and harmlessness. The paper explicitly warns that the resulting models can follow harmful instructions, fabricate facts and produce biased or toxic content. Preference improvement therefore leaves a question central to Preference Aggregation: whose judgments count, and how should conflicts between users, affected people and designers be resolved? The study describes that problem rather than settling it.[2]
Alignment relevance and limits
InstructGPT supplies empirical evidence that demonstration learning plus preference-based reinforcement learning can change broad language-task behavior without another large increase in model size. It also shows why Scalable Oversight depends on what humans can reliably judge: optimizing a favorable assessment is insufficient when mistakes are difficult to detect or the target group omits affected stakeholders. Generalization to tasks beyond the training distribution requires further study; this is not a demonstration of reliable supervision of superhuman systems.[2]
Historical context
The original arXiv v1 was submitted on March 4, 2022. This article uses that supported submission date for its single chronology entry. The paper applies earlier human-feedback methods to a broad instruction distribution; it does not introduce reinforcement learning from human feedback from scratch. Its discussion places the work in an iterative alignment research program focused on existing systems.[2]
A subsequent change in the supervision channel
Constitutional AI: Harmlessness from AI Feedback subsequently tests replacing human harmlessness comparisons with model judgments guided by written principles, while retaining human helpfulness feedback. Read together, the two papers distinguish improving instruction following from choosing and enforcing limits on which requests an assistant should fulfill. Their reported evaluations use different setups and should not be treated as a direct performance comparison.
Evidence scope
Selected original v1 methods, findings and discussion were inspected, with the pipeline on rendered PDF page 3 and the toxicity comparison on page 14 checked visually. The full appendices, cited studies, implementation and independent replication were not audited. The original PDF copy is ready in the research notebook; indexed identity and selected prose were checked, without certifying complete extraction fidelity.[2]
Earlier preference-learning experiments
Deep reinforcement learning from human preferences supplies a 2017 comparison point: a reward model learned from short behavior clips guides agents in games and simulated robotics, with continuing feedback as behavior changes. InstructGPT combines demonstrations and response rankings for broad language tasks. The common pattern is learning a reward from judgments; the supervisory data, policy optimization and evaluation conditions differ.[3]
A narrower language-task predecessor
Learning to summarize from human feedback studies a narrower 2020 task: a learned reward guides summaries of Reddit posts, with transfer to news articles. Its optimization experiment shows predicted reward diverging from human preference. Read alongside InstructGPT, it separates a reusable training pattern from the task-specific judgments that define success.[1]
Sources
- Learning to summarize from human feedback · Source record src-187 · Back to claim ↑1
- Training language models to follow instructions with human feedback · Source record src-179 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7 ↑8 ↑9 ↑10
- Deep reinforcement learning from human preferences · Source record src-185 · Back to claim ↑1
Pages that link here
Last updated 2026-10-10