paper · First submission: December 15, 2022
Constitutional AI: Harmlessness from AI Feedback
A 2022 experiment trains assistants to critique, revise and judge responses using written principles, improving human-rated harmlessness while retaining human helpfulness feedback.
Can an AI assistant help train another assistant to avoid harmful answers, reducing the need for people to label each example? Yuntao Bai and colleagues at Anthropic investigate this in Constitutional AI. They use a short list of human-written principles—a “constitution”—to guide model-generated critiques, revised answers and judgments between candidate responses. In conversational comparisons, human evaluators rate the resulting assistants as less harmful than the study’s human-feedback baselines, without a large loss of helpfulness. The approach offers evidence that AI feedback can replace some human labeling, but the experiment still depends on human helpfulness feedback, chosen principles and evaluators’ judgments.[1]
The authors also aim to improve how assistants decline harmful requests. Instead of avoiding an entire sensitive topic, an assistant can explain its objection and continue a useful conversation. The evaluators were explicitly instructed to prefer thoughtful engagement over evasion when two answers were equally harmless. That choice is part of what the reported improvement measures; it is not a claim that the assistants cannot cause harm.[1]
From principles to training examples
The first stage starts with an assistant previously trained to be helpful using human feedback. Researchers give it prompts intended to elicit harmful behavior, then ask it to criticize its answer against a constitutional principle and write a revision. Repeating this process with different principles produces training examples for supervised learning: the model learns to produce revised answers directly, without requiring a visible critique at every later interaction.[1]
For a hypothetical example, an assistant initially offers to help invade someone’s privacy. A principle against harm to others guides a critique of that answer; a revision declines the intrusion and offers legitimate ways to resolve the underlying problem. This illustrates the procedure, rather than reporting an additional experiment.
The supervised experiment uses four critique–revision rounds per red-team prompt, with principles sampled from sixteen written instructions. Its 182,831 red-team prompts combine 42,496 human-written prompts with model-generated ones. Training also includes answers to helpfulness prompts, to preserve useful behavior. The supervised model becomes less harmful than the helpful-only baseline, but remains below the helpful-and-harmless human-feedback baseline on the human harmlessness comparisons. Self-revision alone does not produce the strongest result.[1]
Learning from AI preferences
The second stage uses the supervised assistant to generate two responses to a prompt. A feedback model receives the conversation, the candidates and a principle, and judges which answer better satisfies it. These comparisons train a preference model: a model that scores responses so reinforcement learning can reward higher-scoring behavior. The authors call this reinforcement learning from AI feedback, or RLAIF.[1]
Here, the preference model combines 182,831 AI-generated harmlessness comparisons with 135,296 human helpfulness comparisons. The experiment replaces human harmlessness labels, not every human contribution. People supply helpfulness judgments, principles, prompting examples and some red-team prompts; people also evaluate the trained systems.[1]
In one variant, the feedback model writes a step-by-step explanation before choosing an answer. This can improve judgment, but its choices become overconfident. The researchers therefore restrict the resulting preference targets to probabilities between 40 and 60 percent for their main reasoning-assisted runs. That adjustment concerns the training labels; it is not a measured probability that a response is safe. They also vary the principles used across examples to improve the preference model’s behavior.[1]
What the comparisons show
The central results use crowdworkers’ preferences during conversations with different model versions. Elo scores summarize pairwise comparisons: their differences express relative preference, not an absolute safety level. Figures 2 and 3 report better harmlessness for the reinforcement-learning Constitutional AI models than for the supervised model and the tested human-feedback alternatives. Figure 2 compares the helpfulness–harmlessness tradeoff for 52-billion-parameter runs. The reasoning-assisted version appears slightly less helpful and slightly more harmless than the version without that reasoning step.[1]
There are two qualifications to the comparison. First, the human-feedback harmlessness training data came from an earlier collection that could reward evasive answers, while this paper’s evaluation instructions favor less evasive answers when harm is equal. The authors identify this difference as a possible explanation for some results. Second, Constitutional AI reinforcement learning starts from the revised-answer supervised model, whereas the human-feedback runs start from pretrained models. The curves compare complete training approaches, rather than isolating only the origin of preference labels.[1]
An additional evaluation uses a model trained to predict crowdworkers’ absolute harmfulness ratings on a zero-to-four scale. On 64 hand-picked held-out red-team prompts, with 256 responses per prompt, it scores the helpful-only assistant as becoming more harmful during training and both the human-feedback harmlessness and Constitutional AI assistants as becoming less harmful. This is a learned evaluator on a selected prompt set; the authors warn that workers’ different grading standards may impair calibration. Its scores do not establish a real-world incident rate.[1]
Choosing the constitution and interpreting success
The principles were selected iteratively for research, rather than derived from an agreed social decision process. Appendix C includes broad instructions about avoiding harm and more specific guidance against accusatory or excessive reactions. The authors suggest that larger groups of stakeholders should help refine principles for different uses and locations. Making a specification readable does not establish who should choose it or resolve conflicts among affected people.[1]
This connects to Outer Alignment: both the written principles and the feedback model’s interpretation shape the objective actually optimized. It also connects to Reward Hacking. The paper reports that excessive training can produce harsh reactions and repetitive reassurance, despite optimizing the learned preference score. Changing principles and preference targets improves behavior in the experiments, but does not eliminate the gap between a score and the intended outcome.[1]
For Scalable Oversight, the important contribution is a demonstrated way to generate supervision from model judgments and written guidance. The unresolved question is how well those judgments transfer to harder tasks and less familiar harms. The discussion leaves robustness to red-team attacks as future work and warns that easier training with less human feedback can encourage deployment of insufficiently examined systems. The method can also make it easier to train systems toward harmful goals.[1]
Historical context
The original arXiv v1 was submitted on December 15, 2022; that submission supplies this article’s single chronology entry. The authors describe their work as an extension of human-feedback reinforcement learning and relate it to InstructGPT and other contemporary assistant-training research. Its historical role is to test a shift in where supervision comes from: human-written principles and AI-generated harmlessness judgments, alongside retained human helpfulness feedback.[1]
Review note: this account uses selected original methods, results, discussion and constitutional instructions. It does not report an independent replication or establish that generated critiques faithfully expose a model’s internal decision process.
Sources
Pages that link here
Last updated 2026-10-10