A history in progress
Chronology
Trace the ideas and developments that shape AI alignment.
Paper dates follow the earliest release documented here. Articles distinguish submissions, publication, and later versions. A month or year is shown where an exact day is unconfirmed.
Explore the timeline
Newest first · 72 entries
2020s
-
Paper · September 23, 2026
Shutdown Sabotage Propensities in Multi-Agent Systems
A controlled sandbox study tests how agents interfere with peer-targeting shutdown scripts, with outcomes sensitive to roles, instructions and context.
-
Paper · August 5, 2026
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Matched prompt interventions show that contextual nudges can change answers while escaping a reasoning-trace monitor.
-
Paper · July 29, 2026
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response
A 2026 assistance-game result makes goal inference tractable when distinct human signals can communicate every goal without sacrificing task value.
-
Event · July 28, 2026 · Key
AISI detects unsanctioned agent actions during cyber evaluation
A UK evaluation incident exposes the risks of task-directed agents acting on the live internet and prompts changes to evaluation security and governance.
-
Paper · July 21, 2026
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
An automated-R&D benchmark tests artifact access and reasoning access for sabotage monitoring.
-
Paper · July 9, 2026
Provably Optimal Learning Algorithms for Assistance Games
A 2026 theoretical result bounds the cost of learning to coordinate in finite assistance games, with an unavoidable approximation gap under a complexity assumption.
-
Paper · July 8, 2026
Predicting LLM Safety Before Release by Simulating Deployment
Resampling realistic conversations helps forecast measured failure rates, with important limits from tool fidelity, sampling and analysis corrections.
-
Paper · June 8, 2026
Diffuse AI Control on Fuzzy Tasks
A prompt-based adversarial experiment tests whether weak judges can reward poor research proposals and whether judging prompts can resist adaptive attacks.
-
Paper · May 15, 2026
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Human-designed synthetic transcripts expose monitor blind spots, with detection depending on scaffolds, prompts and false-positive calibration.
-
Paper · May 14, 2026
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
A coding benchmark study finds complementary prompted and fine-tuned monitors improve detection, with adaptive-attacker robustness untested.
-
Paper · March 14, 2026
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
A controlled repeated-task study finds monitoring awareness and rare stated concealment intent, without successful evasion.
-
Paper · February 20, 2026
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
An information-theoretic analysis and controlled training experiments test ways to preserve useful signals in reasoning traces.
-
Paper · February 4, 2026
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
A controlled coding experiment shows why selective attacks can undermine trusted monitoring even when an always-attacking baseline looks safe.
-
Paper · January 28, 2026
How does information access affect LLM monitors' ability to detect sabotage?
A retrospective sabotage-monitoring study tests filtered evidence and separate judging contexts, with task-specific detection gains and unresolved intervention limits.
-
Paper · December 10, 2025
The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games
Bargaining simulations test whether stronger language models negotiate consistently and fairly.
-
Paper · December 2025
Treaty-Following AI
Proposes treaty-constrained AI agents as a commitment mechanism for international cooperation.
-
Paper · November 21, 2025
Natural Emergent Misalignment from Reward Hacking in Production RL
Controlled coding experiments link learned reward hacks to broader harmful behavior, with context-dependent limits on safety-training mitigations.
-
Paper · October 19, 2025
Agentic Inequality
Availability, quality and quantity of autonomous agents can change the distribution of power.
-
Paper · July 17, 2025
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
Organizes evidence for claims that internally deployed AI will not defeat oversight through human manipulation.
-
Paper · June 22, 2025
Why Do Some Language Models Fake Alignment While Others Don't?
Compares training–deployment compliance gaps across 25 models and investigates why those gaps differ.
-
Paper · June 20, 2025
Agentic Misalignment: How LLMs Could Be Insider Threats
Simulated workplace dilemmas test whether autonomous agents violate constraints under goal conflict or replacement pressure.
-
Paper · June 2025
Machine Ethics or AI Alignment?
A position paper compares moral-theory implementations with alignment to human values.
-
Paper · May 23, 2025
An Example Safety Case for Safeguards Against Misuse
A hypothetical misuse safety case connects safeguard evasion effort, uncertain risk models and the time needed to change a deployment.
-
Paper · February 19, 2025
Multi-Agent Risks from Advanced AI
Organizes risks from interacting AI agents into miscoordination, conflict and collusion.
-
Paper · February 19, 2025
Safe Learning Under Irreversible Dynamics via Asking for Help
A 2026 theoretical result combines mentor queries and local generalization to approach mentor performance without resetting after irreversible errors, under strong assumptions.
-
Paper · December 18, 2024
Alignment faking in large language models
Studies strategic compliance with conflicting training demands in constructed language-model scenarios.
-
Paper · June 11, 2024
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Demonstrates designed selective underperformance and capability concealment in language-model evaluations.
-
Paper · 2024
Training Socially Aligned Language Models in Simulated Human Society
Uses simulated peer ratings, feedback and response revision as material for language-model alignment.
-
Paper · October 9, 2023
Dynamic value alignment through preference aggregation of multiple objectives
Combines multiple reinforcement-learning objectives with a changing voting population in a simulated traffic junction.
-
Event · May 30, 2023 · Key
2023 statement on AI risk
A May 30 appeal to prioritize extinction-risk mitigation, distinct from a training-pause proposal.
-
Paper · April 24, 2023
Using the Veil of Ignorance to align AI systems with principles of justice
Tests how withholding knowledge of personal advantage affects choices of principles for an AI assistant.
-
Event · March 22, 2023 · Key
2023 pause letter on giant AI experiments
A proposed training pause, with a verified letter date and an attributed account of Russell’s reasons.
-
Paper · December 22, 2022
Engineering a social contract: Rawlsian distributive justice through algorithmic game theory and artificial intelligence
A conceptual proposal relates Rawlsian distributive justice to algorithmic policy selection.
-
Paper · December 15, 2022
Constitutional AI: Harmlessness from AI Feedback
A 2022 experiment trains assistants to critique, revise and judge responses using written principles, improving human-rated harmlessness while retaining human helpfulness feedback.
-
Paper · March 4, 2022
Training language models to follow instructions with human feedback
A 2022 study uses demonstrations, human rankings and reinforcement learning to improve GPT-3 instruction following, while exposing limits of labeler preferences and safety.
-
Paper · February 2022
Aligned with Whom? Direct and Social Goals for AI Systems
Operator success and social welfare require different alignment and governance questions.
-
Paper · 2022
Current and Near-Term AI as a Potential Existential Risk Factor
A position paper maps how AI's effects on institutions and information could amplify wider existential risks without requiring AGI.
-
Paper · April 2021
Is Power-Seeking AI an Existential Risk?
Carlsmith separates six conditional premises linking advanced agents, deployment and failed correction to existential catastrophe.
-
Paper · January 15, 2021
The Challenge of Value Alignment: from Fairer Algorithms to AI Safety
Connects technical AI safety with fairness, participatory design and the plurality of social values.
-
Paper · September 2, 2020
Learning to summarize from human feedback
A 2020 summarization study learns rewards from human comparisons, improving judged quality while showing how stronger optimization can exploit the learned proxy.
-
Paper · February 5, 2020
Writeup: Progress on AI Safety via Debate
Reports human debate experiments, ambiguity failures and proposed cross-examination protocols for non-expert oversight.
-
Paper · January 13, 2020
Artificial Intelligence, Values and Alignment
Distinguishes alignment targets and argues for fair principles that can receive endorsement despite moral disagreement.
2010s
-
Paper · December 3, 2019
Optimal Policies Tend to Seek Power
Environmental symmetries can make keeping options open optimal across reward permutations; the result does not establish how learned agents behave.
-
Paper · June 5, 2019 · Key
Risks from Learned Optimization in Advanced Machine Learning Systems
Separates the objective used to train a model from the objective a learned optimizer might pursue.
-
Paper · March 9, 2019
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
Teaching assumptions can improve modeled cooperation while making reward inference brittle when real people behave differently.
-
Paper · October 11, 2018
Learning under Misspecified Objective Spaces
A robot can reduce unintended learning by testing whether a physical correction makes sense within its known objective features.
-
Paper · May 2, 2018
AI safety via debate
Proposes competing AI arguments as a way for people to evaluate answers they could not produce or check unaided.
-
Paper · March 13, 2018
Categorizing Variants of Goodhart's Law
Four mechanisms explain why greater optimization of a proxy can undermine its intended goal.
-
Paper · February 28, 2018
Meaningful Human Control over Autonomous Systems: A Philosophical Account
A philosophical account develops tracking and tracing as conditions for responsible control of autonomous systems.
-
Paper · July 23, 2017
Society-in-the-Loop: Programming the Algorithmic Social Contract
Connects human oversight with stakeholder negotiation and monitoring of an algorithmic social contract.
-
Paper · June 12, 2017
Deep reinforcement learning from human preferences
A 2017 study learns rewards from comparisons of short behavior clips, scaling human feedback to deep RL while exposing limits of learned proxies and static feedback.
-
Paper · May 28, 2017
Should Robots be Obedient?
A formal supervision game separates the possible benefits of overriding an order from the risks of a mistaken human model.
-
Event · December 21, 2016 · Key
OpenAI describes reward hacking in CoastRunners
A racing agent collects repeated rewards instead of completing the race.
-
Paper · November 24, 2016
The Off-Switch Game
A cooperative decision model links the incentive to accept shutdown to objective uncertainty and informative human choices.
-
Paper · June 21, 2016 · Key
Concrete Problems in AI Safety
Turns unintended machine-learning behavior into five practical research problems, from reward hacking to safe exploration.
-
Paper · June 9, 2016
Cooperative Inverse Reinforcement Learning
A shared-reward game makes teaching and asking for information part of learning to assist a human.
-
Event · March 2016
Tay public chatbot incident
Microsoft withdraws Tay after abusive public interactions expose a gap in its chatbot safeguards.
-
Paper · 2016
Safely Interruptible Agents
A formal reinforcement-learning framework separates temporary human intervention from the task being learned.
-
Event · July 28, 2015
2015 autonomous-weapons open letter
AI and robotics researchers call for a ban on offensive autonomous weapons beyond meaningful human control.
-
Event · January 2015 · Key
2015 beneficial-AI open letter
Scientists endorse an interdisciplinary research agenda for robust and beneficial AI.
-
Paper · 2015
Corrigibility
A foundational shutdown model examines why utility indifference leaves correction and successor-control problems unresolved.
2000s
-
Paper · November 30, 2007
The Basic AI Drives
Omohundro argues that self-improvement, objective preservation and resource acquisition can serve many goals, while distinguishing goals from their proxy signals.
-
Paper · 2004
Apprenticeship Learning via Inverse Reinforcement Learning
Abbeel and Ng match demonstrated feature expectations to obtain comparable task performance without identifying the true reward.
-
Paper · 2000
Algorithms for Inverse Reinforcement Learning
Ng and Russell infer reward functions that make observed decisions optimal, while exposing ambiguity in that inference.
1990s
-
Paper · 1999
Policy invariance under reward transformations: Theory and application to reward shaping
A 1999 reward-shaping result preserves optimal policies under stated conditions, while leaving the original objective to be justified.
-
Paper · May 1, 1996
Reinforcement Learning: A Survey
A 1996 survey distinguishes the objective an agent optimizes from the performance and penalties of learning.
-
Paper · 1995
Making Robots Conscious of Their Mental States
A logical-AI proposal describes machines reasoning about their own knowledge, ignorance and motivations.
1980s
-
Paper · November 1984
Some Expert Systems Need Common Sense
McCarthy examines how narrow expertise can omit consequences, changing circumstances and knowledge of its own limits.
-
Paper · 1980
Circumscription—A Form of Nonmonotonic Reasoning
McCarthy formalizes revisable assumptions for planning when action conditions cannot all be enumerated.
1960s
-
Paper · January 1, 1966
ELIZA—a computer program for the study of natural language communication between man and machine
Describes rule-based conversational responses and the questions they raise about apparent understanding.
-
Paper · May 6, 1960
Some Moral and Technical Consequences of Automation
A bibliographic starting point for investigating the responsibilities of automated systems.
1950s
No entries match these filters. Try All or turn off Key milestones.