note

Key Papers: A Reading Guide

Question-led reading sequences through alignment research, with reasons for each step and prompts for assessing the evidence.

These papers offer different ways into the alignment conversation. This guide is a curated starting point, rather than a comprehensive bibliography or a ranking. Each linked writeup includes original sources, the paper’s contribution, and questions it leaves open.

If you are new to alignment

Read the introduction to AI alignment first. Then take this short route:

  1. Start with Reward Hacking. In its example, identify the intended task, the rewarded behavior and the action that separates them. This gives you a concrete question to carry into the papers: what does a successful score actually establish?
  2. Read Concrete Problems in AI Safety, beginning with the overview and its division of problems. For each example, ask whether the trouble enters through the objective, the available supervision or changed conditions. This step turns one failure into a way to distinguish problems that may need different remedies.
  3. Read Artificial Intelligence, Values and Alignment. Ask whose instructions, interests or values would supply the target in your earlier example, and whether those targets conflict. This follows the technical agenda because choosing a better score still leaves the purpose of that score to be justified.

This sequence moves from an example to a research agenda to the choice of alignment target. It is an editorial reading route, not a claim that these papers exhaust the field or that research developed in a single line. If you already know the basics, choose a question below.

What should a system be aligned with?

Start with Artificial Intelligence, Values and Alignment for the distinctions between instructions, interests, and values. Aligned with Whom? Direct and Social Goals for AI Systems connects the operator’s objectives to costs borne by other people. Together they introduce why Direct Alignment and Social Alignment need separate attention.

How can training objectives go wrong?

Use this sequence when you want to locate a failure, rather than give every unwanted outcome the same name:

  1. Locate the problem: Concrete Problems in AI Safety. Choose one example and write down what a proposed defense would have to preserve. Then check the writeup’s distinction between a suggested research direction and a demonstrated result. You can skip this step if you just completed the introductory route.
  2. Examine the proxy: Categorizing Variants of Goodhart’s Law. Keep that example in mind while reading the distinction between selection and intervention. Ask what evidence would distinguish noise, a relationship failing outside its familiar range, a changed causal system and another agent responding to incentives. This follows the agenda by asking why optimizing a measurement can disappoint, rather than merely observing that it does.[5]
  3. Examine the learned mechanism: Risks from Learned Optimization in Advanced Machine Learning Systems, especially sections 1.1–1.3 and 3.1. Track the training objective, the criterion used by any internal search, and the objective an observer infers from behavior. This follows the proxy question by asking where a different criterion might enter the learned system itself. The framework requires internal optimization; an unwanted action alone does not establish it.[2]

Pause before the deceptive case in section 4. Can you explain an objective mismatch without assuming that the system knows it is being trained? Inner Alignment and Mesa-Optimization clarify the terms. Continue to the behavioral studies below when you want to compare a proposed mechanism with experimental evidence, keeping that distinction open.

How can people oversee difficult work?

AI safety via debate proposes adversarial arguments judged by people. Training Socially Aligned Language Models in Simulated Human Society explores simulated social feedback as a training resource. Read them alongside Scalable Oversight: both move beyond checking each answer directly, while raising questions about the reliability of the feedback.

What do demonstrations of deceptive behavior establish?

AI Sandbagging: Language Models can Strategically Underperform on Evaluations concerns concealed capabilities. Alignment faking in large language models concerns strategic compliance during training. Agentic Misalignment: How LLMs Could Be Insider Threats examines harmful actions under workplace pressures, and Natural Emergent Misalignment from Reward Hacking in Production RL examines generalization from learned exploits.

Read their experimental conditions carefully. These demonstrations address different behaviors; their writeups distinguish observed results from claims about prevalence in ordinary deployment.

For a focused comparison of feedback and representation, see Who Supplies Alignment Feedback?.

How should evaluation evidence be read?

Compare the Safety Case entry with the experiments above. Ask what was tested, what the authors conclude, and which assumptions are needed to carry a result into deployment. A formal theorem, a controlled behavioral study, and a philosophical argument require different kinds of scrutiny. Follow each writeup to its original sources and stated review scope.

How does alignment extend to collective life?

Society-in-the-Loop: Programming the Algorithmic Social Contract proposes negotiation and monitoring among affected people. Agentic Inequality asks how access to agents distributes power. Treaty-Following AI explores constraints that could support international commitments. They extend the conversation toward institutions, distribution, and enforceable cooperation.

Use the chronology to follow the development of these questions, or browse the Sources for original texts and related records.

Sources

  1. Concrete Problems in AI Safety · Source record ref-cd3035dbef6c
  2. Risks from Learned Optimization in Advanced Machine Learning Systems · Source record ref-c4858d4ef280 · Back to claim ↑1
  3. Artificial Intelligence, Values and Alignment · Source record ref-030a9cc7079a
  4. Society-in-the-Loop: Programming the Algorithmic Social Contract · Source record ref-ad43ed014e22
  5. Categorizing Variants of Goodhart's Law · Source record ref-e178cbf073c0 · Back to claim ↑1
  6. AI safety via debate · Source record ref-debate-2018
  7. Aligned with Whom? Direct and Social Goals for AI Systems · Source record src-016
  8. TRAINING SOCIALLY ALIGNED LANGUAGE MODELS IN SIMULATED HUMAN SOCIETY · Source record src-050
  9. AI Sandbagging: Language Models can Strategically Underperform on Evaluations · Source record src-011
  10. ALIGNMENT FAKING IN LARGE LANGUAGE MODELS · Source record src-012
  11. Agentic Misalignment: How LLMs Could Be Insider Threats · Source record src-015
  12. Natural emergent misalignment from reward hacking in production RL · Source record src-037
  13. Agentic Inequality · Source record src-013
  14. Treaty-Following AI · Source record src-071

Pages that link here

Last updated 2026-10-11