note

Origins of Alignment: Problems Before the Name

A historical reading route separates intelligent performance, communicating purposes, human judgment and the limits of delegation.

Which older problems help us understand alignment today, without treating every early AI paper as alignment research? Start with recognition, move to purpose and limits, then separate capability advances from work on supervision. This is a selective reading route, not a genealogy proving who influenced whom.

On this page: Recognition and purpose · Delegation · Learning theory · Modern branches · Gaps

Recognition, learning and purpose

  1. 1950: Computing Machinery and Intelligence. Read sections 1 and 7 together. The imitation game asks about recognizable performance; educating a child-machine raises questions about teaching and communication. Ask what a reward signal conveys and what it leaves unsaid.[1]
  2. 1960: Some Moral and Technical Consequences of Automation. Move from teaching to the consequences of acting. Look for the distinction between a communicated purpose and the desired outcome, then ask whether people can intervene before an action becomes irreversible. Wiener supplies an argument and examples, not a modern loss-of-control experiment.[2]
  3. 1966: ELIZA. Read the original program description, then compare the ELIZA Effect. Ask what the mechanism does and what a user takes the conversation to mean. This returns to recognition with a practical caution: an impression of understanding and dependable assistance need different evidence.[3]

These questions remain useful even when the architectures change. They do not establish that today’s systems reproduce an early program’s mechanism.

Expertise and the limits of delegation

  1. 1984: Some Expert Systems Need Common Sense. Track what an adviser represents about treatment, consequences and the physician’s role. McCarthy’s examples ask whether narrow competence leaves important context out. His discussion also allows that a physician may supply the missing common sense; do not turn it into a claim that every useful system needs general intelligence.[4]
  2. 1995: Making Robots Conscious of Their Mental States. Follow that limitation into a proposal for representing knowledge and ignorance. Ask what a robot could do differently if it recognized that it lacked information. Logical self-knowledge is a proposed functional ability, not proof of subjective experience or safe deployment.[5]

The transition is from the limits of a task model to recognizing those limits. John McCarthy connects these entries; Meaningful Human Control asks the later responsibility question of who can understand and act on them.

Learning theory and public risk

  1. 1996–1999: Reinforcement Learning: A Survey and Policy invariance under reward transformations: Theory and application to reward shaping. First identify the learning objective and performance criterion. Then ask which transformations preserve that objective. Preserving an optimal policy under a formal model leaves the desirability of the original reward unresolved.[6][7]
  2. 2000: Bill Joy’s Why the Future Doesn’t Need Us. Compare the technical question with a public argument about relinquishing dangerous technological capabilities. Joy discusses robotics alongside genetics and nanotechnology, including self-replication. Read it as a technologist’s warning and proposed response, not an empirical demonstration or a prediction already vindicated.[8]

These readings juxtapose formal guarantees and public choices about technological development. They do not establish a causal path from one to the other.

Two branches into modern AI

  1. 2015–2016: Research Priorities for Robust and Beneficial Artificial Intelligence → Concrete Problems in AI Safety. Ask how verification, validity, security and control divide the task, then how the accident-risk agenda turns concerns into research problems. An agenda identifies work to do; it does not certify solutions.[9]
  2. Capability branch: Attention Is All You Need. Identify the computational constraint the Transformer changes and the tasks used to evaluate it. The paper credits eight collaborators. Architecture and task performance leave goals and governance to be addressed separately.[10]
  3. Supervision branch: Deep reinforcement learning from human preferences → Learning to summarize from human feedback. Trace human comparisons through a learned reward into trained behavior. Ask when predicting judgments and satisfying people come apart.[11][12] Continue through the oversight sequence in Key Papers: A Reading Guide, the central home for reading journeys.

Geoffrey Hinton provides one connection between earlier neural-network research and later public risk advocacy. Demis Hassabis offers an institutional research perspective. Neither biography substitutes for evaluating the underlying research.

A governance branch asks who can report risk information. The June 4, 2024 Right to Warn letter, signed by Daniel Kokotajlo and others, proposes reporting channels and protections against retaliation. Compare that institutional proposal with the technical supervision methods above; a way to report concerns is not a validated model-control method.[13]

What this history still needs

The 1970s are a substantive gap, especially Weizenbaum’s arguments about which judgments should be delegated. The Dartmouth proposal, I. J. Good’s intelligence-explosion argument, earlier control theory and 2000s safety proposals also need original-passage reviews before fuller treatment. This guide makes no claim about the first use of “alignment,” one founding figure, or an uninterrupted intellectual lineage.

What evidence would demonstrate influence: citations, correspondence, interviews or later recollections? How did researchers’ proposed remedies change as learning replaced hand-written rules? Which early warnings concerned capability, miscommunication, institutional power or moral authority? These are different historical questions.

Review scope: selected Turing and Wiener originals newly checked; the intermediate scholarly readings rely on the scopes and limitations in their existing writeups. Selected Joy passages and Transformer v1 introduction/attribution checked; the Right to Warn letter body and selected names reviewed. No comprehensive primary-literature or reception review is claimed.

Sources

  1. Alan Turing, Computing Machinery and Intelligence (Mind, October 1950) · Source record ref-c0af2eabf7fd · Back to claim ↑1
  2. Norbert Wiener, Some Moral and Technical Consequences of Automation (Science, 1960) · Source record ref-226167410ec0 · Back to claim ↑1
  3. Joseph Weizenbaum, ELIZA (Communications of the ACM, January 1966) · Source record ref-cddb0c54767c · Back to claim ↑1
  4. Some Expert Systems Need Common Sense · Source record src-042 · Back to claim ↑1
  5. Making Robots Conscious of Their Mental States · Source record src-032 · Back to claim ↑1
  6. Reinforcement Learning: A Survey · Source record src-176 · Back to claim ↑1
  7. Policy invariance under reward transformations: Theory and application to reward shaping · Source record src-177 · Back to claim ↑1
  8. Why the Future Doesn’t Need Us · Source record ref-joy-2000 · Back to claim ↑1
  9. Research Priorities for Robust and Beneficial Artificial Intelligence · Source record src-191 · Back to claim ↑1
  10. Attention Is All You Need · Source record ref-transformer-2017 · Back to claim ↑1
  11. Deep reinforcement learning from human preferences · Source record src-185 · Back to claim ↑1
  12. Learning to summarize from human feedback · Source record src-187 · Back to claim ↑1
  13. A Right to Warn about Advanced Artificial Intelligence · Source record ref-right-to-warn-2024 · Back to claim ↑1

Last updated 2026-10-11