person
Stuart Russell
AI researcher whose work represented here connects reward learning and a beneficial-AI research agenda.
Stuart Russell coauthored Research Priorities for Robust and Beneficial Artificial Intelligence with Daniel Dewey and Max Tegmark. The journal identifies his affiliation as the University of California, Berkeley. The 2015 beneficial-AI open letter lists him as a prominent signatory.[3][4]
His technical work represented here includes Algorithms for Inverse Reinforcement Learning with Andrew Ng and Policy invariance under reward transformations: Theory and application to reward shaping with Ng and Daishi Harada.[5][6] These contributions address inferring rewards and preserving optimal policies under reward transformations. They do not themselves establish that the resulting objective is socially desirable.
Objectives as a research responsibility
In Science Friday’s September 23, 2016 interview, Russell argued that choosing objectives should be part of AI research. Its published transcript states: “getting the right objectives actually becomes part of the job of the field” (answer following Flatow’s question about why a human-compatibility center is needed). The program introduced him as leading Berkeley’s new Center for Human Compatible Artificial Intelligence. This documents his research aim, not an achieved safety guarantee; the quotation is from the transcript, without an audio cross-check.[7]
Learning with irreversible consequences
Russell also coauthored Safe Learning Under Irreversible Dynamics via Asking for Help with Benjamin Plaut, Juan Liévano-Karim and Hanlin Zhu. The JMLR 2026 paper studies agents that learn without resets and can ask a mentor before acting. The authors summarize the motivation in section 1.1: “finding a good policy is futile if irreparable damage is caused along the way” (p.2). This connects their technical learning criterion to the cost of mistakes during learning. The result remains mentor-relative and assumption-dependent; section 7 says the algorithm is not ready for practical use.[2]
Autonomous-weapons advocacy
FLI’s July 29, 2015 report identifies Russell and Toby Walsh as announcing the 2015 autonomous-weapons open letter at an IJCAI press conference in Buenos Aires. This documents his role in that appeal for restrictions on military autonomy, beyond his technical work on reward learning. The announcement is an institutional account, rather than an independent assessment of the campaign’s effects.[8]
Cooperative assistance and correction
With Dylan Hadfield-Menell, Anca Dragan and Pieter Abbeel, Russell coauthored Cooperative Inverse Reinforcement Learning and The Off-Switch Game. The first models teaching and learning as joint pursuit of a human reward; the second studies when uncertainty and informative human choices make oversight valuable. These formal contributions connect his objective-selection research aim to specific incentive models, while leaving their human-model and computation assumptions in place.[9][10]
Russell also coauthored Should Robots be Obedient?. That paper examines how incorrect assumptions about human preferences and behavior affect the modeled advantage of autonomy. Its formal tradeoff is not an unrestricted recommendation to override people.[1]
Public risk appeals in 2023
Russell appears in the selected signatories of the 2023 pause letter on giant AI experiments and in the original announcement of the 2023 statement on AI risk. The first proposes a training pause; the second prioritizes extinction-risk mitigation without prescribing that intervention.[11][13]
In Science Friday’s April 7 interview, he qualified his discussion of future loss of control: “the current systems do not present that risk, as far as we know” (answer to Flatow’s question invoking Stephen Hawking). This preserves the distinction between his immediate safety concerns and his forecasts about stronger future systems. The quotation is checked against the published transcript; its audio has not been cross-checked.[14]
Sources
- Should Robots be Obedient? · Source record src-208 · Back to claim ↑1
- Safe Learning Under Irreversible Dynamics via Asking for Help · Source record src-178 · Back to claim ↑1
- Research Priorities for Robust and Beneficial Artificial Intelligence · Source record src-191 · Back to claim ↑1
- Research Priorities for Robust and Beneficial Artificial Intelligence: An Open Letter · Source record src-190 · Back to claim ↑1
- Algorithms for Inverse Reinforcement Learning · Source record src-172 · Back to claim ↑1
- Policy invariance under reward transformations: Theory and application to reward shaping · Source record src-177 · Back to claim ↑1
- Making the Most of A.I.’s Potential · Source record src-194 · Back to claim ↑1
- Open letter on AI weapons · Source record src-200 · Back to claim ↑1
- The Off-Switch Game · Source record src-203 · Back to claim ↑1
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1
- Pause Giant AI Experiments: An Open Letter · Source record src-210 · Back to claim ↑1
- Statement on AI Risk · Source record src-211
- AI Extinction Statement Press Release · Source record src-212 · Back to claim ↑1
- An Open Letter Asks AI Researchers To Reconsider Responsibilities · Source record src-213 · Back to claim ↑1
Pages that link here
- 2015 autonomous-weapons open letter event
- 2015 beneficial-AI open letter event
- 2023 pause letter on giant AI experiments event
- 2023 statement on AI risk event
- Algorithms for Inverse Reinforcement Learning paper
- Cooperative Inverse Reinforcement Learning paper
- Provably Optimal Learning Algorithms for Assistance Games paper
- Research Priorities for Robust and Beneficial Artificial Intelligence paper
- Should Robots be Obedient? paper
- The Off-Switch Game paper
Last updated 2026-10-10