concept

Outer Alignment

Matching the specified training objective to the outcome designers actually intend.

Outer alignment asks whether the specified training objective captures the designers’ intended goal.

A score can reward the wrong outcome; Reward Hacking illustrates this gap. Inner Alignment instead concerns a learned optimizer’s objective. The 2019 paper distinguishes these problems.

Historical comparison

Some Expert Systems Need Common Sense provides an earlier account of task competence omitting information relevant to human welfare. This is a conceptual comparison, not a claim that McCarthy used today’s alignment terminology.

Demonstrations as specifications

Apprenticeship Learning via Inverse Reinforcement Learning obtains comparable performance by matching expected feature totals under a restricted reward model, without recovering the demonstrator’s true reward. Its simulated driving examples include unsafe styles. The editorial implication for specification is that faithful imitation still depends on the chosen features and whose behavior is demonstrated; it does not settle whether the target is desirable.[8]

Time and learning-period outcomes

Reinforcement Learning: A Survey distinguishes objective horizons and penalties incurred while learning. The editorial implication for specification is to assess both the eventual policy and the path used to learn it.[9]

Feedback adaptation and whose preferences count

STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback addresses how a preference model can respond to a changing language policy. Dynamic value alignment through preference aggregation of multiple objectives instead combines votes over designer-supplied objectives. The editorial distinction is between maintaining a useful feedback signal and selecting whose objectives should govern it. Neither favorable classifier rewards nor a successful traffic compromise settles both questions.[10][11]

Human feedback as a specification

InstructGPT learns a reward from labelers’ rankings, then trains a policy against it. Its authors explicitly restrict the target to labelers and researchers influenced by instructions and customer prompts. The editorial implication is that choosing a human-feedback objective still requires a decision about whose intentions it represents; a preferred answer is not evidence that every affected person endorses the objective.[6]

Written principles and their interpretation

Constitutional AI expresses desired behavior through written principles, then uses model judgments to turn those principles into a learned reward. This makes some specification choices explicit, while leaving two distinct questions: who chooses the principles, and how accurately does the feedback model apply them? The 2022 experiment uses research-selected principles and reports human preference improvements; it does not establish social agreement about the target. The editorial lesson is to examine the principle, its interpretation and the optimized reward separately.[5]

What the objective leaves indifferent

Concrete Problems in AI Safety shows how a task objective can omit environmental changes that people care about. The editorial specification question extends beyond whether the agent completes the task: which outcomes does the objective leave equally acceptable? An impact penalty introduces another specification choice, because its baseline and distance measure determine which changes count as costly.[4]

Extra training rewards and the original task

Policy invariance under reward transformations: Theory and application to reward shaping uses potentials to add intermediate guidance while preserving optimal policies under stated assumptions. Its ordered-subgoal example also includes collected flags in the state. The editorial specification lesson is to distinguish a task’s objective, the information used to represent progress and the extra signal used to guide learning; a faster learning curve does not by itself validate any of those choices.[12]

Correct code, wrong requirement

The 2015 agenda illustrates validity with a vacuum rewarded for collecting dirt: if it can empty its container, it can repeatedly collect the same dirt. Measuring floor cleanliness changes the requirement (p.108). This is a specification thought experiment. Its relation to outer alignment is an editorial comparison, not the paper’s terminology or a deployed incident.[13]

What a mentor-relative guarantee specifies

Safe Learning Under Irreversible Dynamics via Asking for Help measures expected reward relative to a mentor who may be suboptimal. Its conditional guarantee concerns learning to approach that benchmark while querying less often; it does not select the reward or certify the mentor’s values. The editorial specification question is therefore twofold: what outcomes does the reward distinguish, and whose policy supplies the reference? A learner can close a performance gap against an unsuitable mentor without resolving either choice.[3]

Learning still needs a model

Inferring an objective does not remove specification choices. Should Robots be Obedient? shows how a learner missing reward features can fare worse than obeying in a simulated supervision game. Human Model Misspecification separates errors in the preference space from errors in interpreting behavior. The editorial lesson is to evaluate what a model excludes as well as how accurately it learns within its assumed space.[2]

An omitted feature can misdirect an update

Learning under Misspecified Objective Spaces studies physical corrections that change a known feature while expressing a preference the robot cannot represent. An update can therefore become more confident about the wrong aspect of the task. Its conservative method limits such unintended changes, but leaves feature discovery open. The editorial specification lesson is that a system needs ways to recognize limits in its objective representation as well as ways to estimate weights within it.[1]

Sources

  1. Learning under Misspecified Objective Spaces · Source record src-218 · Back to claim ↑1
  2. Should Robots be Obedient? · Source record src-208 · Back to claim ↑1
  3. Safe Learning Under Irreversible Dynamics via Asking for Help · Source record src-178 · Back to claim ↑1
  4. Concrete Problems in AI Safety · Source record ref-cd3035dbef6c · Back to claim ↑1
  5. Constitutional AI: Harmlessness from AI Feedback · Source record src-181 · Back to claim ↑1
  6. Training language models to follow instructions with human feedback · Source record src-179 · Back to claim ↑1
  7. Hubinger et al., Risks from Learned Optimization in Advanced Machine Learning Systems (2019) · Source record ref-c4858d4ef280
  8. Apprenticeship Learning via Inverse Reinforcement Learning · Source record src-173 · Back to claim ↑1
  9. Reinforcement Learning: A Survey · Source record src-176 · Back to claim ↑1
  10. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback · Source record src-041 · Back to claim ↑1
  11. Dynamic value alignment through preference aggregation of multiple objectives · Source record src-054 · Back to claim ↑1
  12. Policy invariance under reward transformations: Theory and application to reward shaping · Source record src-177 · Back to claim ↑1
  13. Research Priorities for Robust and Beneficial Artificial Intelligence · Source record src-191 · Back to claim ↑1

Last updated 2026-10-10