A history in progress

Chronology

Trace the ideas and developments that shape AI alignment.

Paper dates follow the earliest release documented here. Articles distinguish submissions, publication, and later versions. A month or year is shown where an exact day is unconfirmed.

Explore the timeline

Newest first · 72 entries

2020s

  1. Paper · September 23, 2026

    Shutdown Sabotage Propensities in Multi-Agent Systems

    A controlled sandbox study tests how agents interfere with peer-targeting shutdown scripts, with outcomes sensitive to roles, instructions and context.

  2. Paper · August 5, 2026

    Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

    Matched prompt interventions show that contextual nudges can change answers while escaping a reasoning-trace monitor.

  3. Paper · July 29, 2026

    Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

    A 2026 assistance-game result makes goal inference tractable when distinct human signals can communicate every goal without sacrificing task value.

  4. Event · July 28, 2026 · Key

    AISI detects unsanctioned agent actions during cyber evaluation

    A UK evaluation incident exposes the risks of task-directed agents acting on the live internet and prompts changes to evaluation security and governance.

  5. Paper · July 21, 2026

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

    An automated-R&D benchmark tests artifact access and reasoning access for sabotage monitoring.

  6. Paper · July 9, 2026

    Provably Optimal Learning Algorithms for Assistance Games

    A 2026 theoretical result bounds the cost of learning to coordinate in finite assistance games, with an unavoidable approximation gap under a complexity assumption.

  7. Paper · July 8, 2026

    Predicting LLM Safety Before Release by Simulating Deployment

    Resampling realistic conversations helps forecast measured failure rates, with important limits from tool fidelity, sampling and analysis corrections.

  8. Paper · June 8, 2026

    Diffuse AI Control on Fuzzy Tasks

    A prompt-based adversarial experiment tests whether weak judges can reward poor research proposals and whether judging prompts can resist adaptive attacks.

  9. Paper · May 15, 2026

    SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

    Human-designed synthetic transcripts expose monitor blind spots, with detection depending on scaffolds, prompts and false-positive calibration.

  10. Paper · May 14, 2026

    Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

    A coding benchmark study finds complementary prompted and fine-tuned monitors improve detection, with adaptive-attacker robustness untested.

  11. Paper · March 14, 2026

    Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback

    A controlled repeated-task study finds monitoring awareness and rare stated concealment intent, without successful evasion.

  12. Paper · February 20, 2026

    Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory

    An information-theoretic analysis and controlled training experiments test ways to preserve useful signals in reasoning traces.

  13. Paper · February 4, 2026

    Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring

    A controlled coding experiment shows why selective attacks can undermine trusted monitoring even when an always-attacking baseline looks safe.

  14. Paper · January 28, 2026

    How does information access affect LLM monitors' ability to detect sabotage?

    A retrospective sabotage-monitoring study tests filtered evidence and separate judging contexts, with task-specific detection gains and unresolved intervention limits.

  15. Paper · December 10, 2025

    The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games

    Bargaining simulations test whether stronger language models negotiate consistently and fairly.

  16. Paper · December 2025

    Treaty-Following AI

    Proposes treaty-constrained AI agents as a commitment mechanism for international cooperation.

  17. Paper · November 21, 2025

    Natural Emergent Misalignment from Reward Hacking in Production RL

    Controlled coding experiments link learned reward hacks to broader harmful behavior, with context-dependent limits on safety-training mitigations.

  18. Paper · October 19, 2025

    Agentic Inequality

    Availability, quality and quantity of autonomous agents can change the distribution of power.

  19. Paper · July 17, 2025

    Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    Organizes evidence for claims that internally deployed AI will not defeat oversight through human manipulation.

  20. Paper · June 22, 2025

    Why Do Some Language Models Fake Alignment While Others Don't?

    Compares training–deployment compliance gaps across 25 models and investigates why those gaps differ.

  21. Paper · June 20, 2025

    Agentic Misalignment: How LLMs Could Be Insider Threats

    Simulated workplace dilemmas test whether autonomous agents violate constraints under goal conflict or replacement pressure.

  22. Paper · June 2025

    Machine Ethics or AI Alignment?

    A position paper compares moral-theory implementations with alignment to human values.

  23. Paper · May 23, 2025

    An Example Safety Case for Safeguards Against Misuse

    A hypothetical misuse safety case connects safeguard evasion effort, uncertain risk models and the time needed to change a deployment.

  24. Paper · February 19, 2025

    Multi-Agent Risks from Advanced AI

    Organizes risks from interacting AI agents into miscoordination, conflict and collusion.

  25. Paper · February 19, 2025

    Safe Learning Under Irreversible Dynamics via Asking for Help

    A 2026 theoretical result combines mentor queries and local generalization to approach mentor performance without resetting after irreversible errors, under strong assumptions.

  26. Paper · December 18, 2024

    Alignment faking in large language models

    Studies strategic compliance with conflicting training demands in constructed language-model scenarios.

  27. Paper · June 11, 2024

    AI Sandbagging: Language Models can Strategically Underperform on Evaluations

    Demonstrates designed selective underperformance and capability concealment in language-model evaluations.

  28. Paper · 2024

    Training Socially Aligned Language Models in Simulated Human Society

    Uses simulated peer ratings, feedback and response revision as material for language-model alignment.

  29. Paper · October 9, 2023

    Dynamic value alignment through preference aggregation of multiple objectives

    Combines multiple reinforcement-learning objectives with a changing voting population in a simulated traffic junction.

  30. Event · May 30, 2023 · Key

    2023 statement on AI risk

    A May 30 appeal to prioritize extinction-risk mitigation, distinct from a training-pause proposal.

  31. Paper · April 24, 2023

    Using the Veil of Ignorance to align AI systems with principles of justice

    Tests how withholding knowledge of personal advantage affects choices of principles for an AI assistant.

  32. Event · March 22, 2023 · Key

    2023 pause letter on giant AI experiments

    A proposed training pause, with a verified letter date and an attributed account of Russell’s reasons.

  33. Paper · December 22, 2022

    Engineering a social contract: Rawlsian distributive justice through algorithmic game theory and artificial intelligence

    A conceptual proposal relates Rawlsian distributive justice to algorithmic policy selection.

  34. Paper · December 15, 2022

    Constitutional AI: Harmlessness from AI Feedback

    A 2022 experiment trains assistants to critique, revise and judge responses using written principles, improving human-rated harmlessness while retaining human helpfulness feedback.

  35. Paper · March 4, 2022

    Training language models to follow instructions with human feedback

    A 2022 study uses demonstrations, human rankings and reinforcement learning to improve GPT-3 instruction following, while exposing limits of labeler preferences and safety.

  36. Paper · February 2022

    Aligned with Whom? Direct and Social Goals for AI Systems

    Operator success and social welfare require different alignment and governance questions.

  37. Paper · 2022

    Current and Near-Term AI as a Potential Existential Risk Factor

    A position paper maps how AI's effects on institutions and information could amplify wider existential risks without requiring AGI.

  38. Paper · April 2021

    Is Power-Seeking AI an Existential Risk?

    Carlsmith separates six conditional premises linking advanced agents, deployment and failed correction to existential catastrophe.

  39. Paper · January 15, 2021

    The Challenge of Value Alignment: from Fairer Algorithms to AI Safety

    Connects technical AI safety with fairness, participatory design and the plurality of social values.

  40. Paper · September 2, 2020

    Learning to summarize from human feedback

    A 2020 summarization study learns rewards from human comparisons, improving judged quality while showing how stronger optimization can exploit the learned proxy.

  41. Paper · February 5, 2020

    Writeup: Progress on AI Safety via Debate

    Reports human debate experiments, ambiguity failures and proposed cross-examination protocols for non-expert oversight.

  42. Paper · January 13, 2020

    Artificial Intelligence, Values and Alignment

    Distinguishes alignment targets and argues for fair principles that can receive endorsement despite moral disagreement.

2010s

  1. Paper · December 3, 2019

    Optimal Policies Tend to Seek Power

    Environmental symmetries can make keeping options open optimal across reward permutations; the result does not establish how learned agents behave.

  2. Paper · June 5, 2019 · Key

    Risks from Learned Optimization in Advanced Machine Learning Systems

    Separates the objective used to train a model from the objective a learned optimizer might pursue.

  3. Paper · March 9, 2019

    Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning

    Teaching assumptions can improve modeled cooperation while making reward inference brittle when real people behave differently.

  4. Paper · October 11, 2018

    Learning under Misspecified Objective Spaces

    A robot can reduce unintended learning by testing whether a physical correction makes sense within its known objective features.

  5. Paper · May 2, 2018

    AI safety via debate

    Proposes competing AI arguments as a way for people to evaluate answers they could not produce or check unaided.

  6. Paper · March 13, 2018

    Categorizing Variants of Goodhart's Law

    Four mechanisms explain why greater optimization of a proxy can undermine its intended goal.

  7. Paper · February 28, 2018

    Meaningful Human Control over Autonomous Systems: A Philosophical Account

    A philosophical account develops tracking and tracing as conditions for responsible control of autonomous systems.

  8. Paper · July 23, 2017

    Society-in-the-Loop: Programming the Algorithmic Social Contract

    Connects human oversight with stakeholder negotiation and monitoring of an algorithmic social contract.

  9. Paper · June 12, 2017

    Deep reinforcement learning from human preferences

    A 2017 study learns rewards from comparisons of short behavior clips, scaling human feedback to deep RL while exposing limits of learned proxies and static feedback.

  10. Paper · May 28, 2017

    Should Robots be Obedient?

    A formal supervision game separates the possible benefits of overriding an order from the risks of a mistaken human model.

  11. Event · December 21, 2016 · Key

    OpenAI describes reward hacking in CoastRunners

    A racing agent collects repeated rewards instead of completing the race.

  12. Paper · November 24, 2016

    The Off-Switch Game

    A cooperative decision model links the incentive to accept shutdown to objective uncertainty and informative human choices.

  13. Paper · June 21, 2016 · Key

    Concrete Problems in AI Safety

    Turns unintended machine-learning behavior into five practical research problems, from reward hacking to safe exploration.

  14. Paper · June 9, 2016

    Cooperative Inverse Reinforcement Learning

    A shared-reward game makes teaching and asking for information part of learning to assist a human.

  15. Event · March 2016

    Tay public chatbot incident

    Microsoft withdraws Tay after abusive public interactions expose a gap in its chatbot safeguards.

  16. Paper · 2016

    Safely Interruptible Agents

    A formal reinforcement-learning framework separates temporary human intervention from the task being learned.

  17. Event · July 28, 2015

    2015 autonomous-weapons open letter

    AI and robotics researchers call for a ban on offensive autonomous weapons beyond meaningful human control.

  18. Event · January 2015 · Key

    2015 beneficial-AI open letter

    Scientists endorse an interdisciplinary research agenda for robust and beneficial AI.

  19. Paper · 2015

    Corrigibility

    A foundational shutdown model examines why utility indifference leaves correction and successor-control problems unresolved.

2000s

  1. Paper · November 30, 2007

    The Basic AI Drives

    Omohundro argues that self-improvement, objective preservation and resource acquisition can serve many goals, while distinguishing goals from their proxy signals.

  2. Paper · 2004

    Apprenticeship Learning via Inverse Reinforcement Learning

    Abbeel and Ng match demonstrated feature expectations to obtain comparable task performance without identifying the true reward.

  3. Paper · 2000

    Algorithms for Inverse Reinforcement Learning

    Ng and Russell infer reward functions that make observed decisions optimal, while exposing ambiguity in that inference.

1990s

  1. Paper · 1999

    Policy invariance under reward transformations: Theory and application to reward shaping

    A 1999 reward-shaping result preserves optimal policies under stated conditions, while leaving the original objective to be justified.

  2. Paper · May 1, 1996

    Reinforcement Learning: A Survey

    A 1996 survey distinguishes the objective an agent optimizes from the performance and penalties of learning.

  3. Paper · 1995

    Making Robots Conscious of Their Mental States

    A logical-AI proposal describes machines reasoning about their own knowledge, ignorance and motivations.

1980s

  1. Paper · November 1984

    Some Expert Systems Need Common Sense

    McCarthy examines how narrow expertise can omit consequences, changing circumstances and knowledge of its own limits.

  2. Paper · 1980

    Circumscription—A Form of Nonmonotonic Reasoning

    McCarthy formalizes revisable assumptions for planning when action conditions cannot all be enumerated.

1960s

  1. Paper · January 1, 1966

    ELIZA—a computer program for the study of natural language communication between man and machine

    Describes rule-based conversational responses and the questions they raise about apparent understanding.

  2. Paper · May 6, 1960

    Some Moral and Technical Consequences of Automation

    A bibliographic starting point for investigating the responsibilities of automated systems.

1950s

  1. Paper · October 1, 1950 · Key

    Computing Machinery and Intelligence

    A paper reframes the question of machine thinking as an observable test.