paper · Research release: April 2021

Is Power-Seeking AI an Existential Risk?

Carlsmith separates six conditional premises linking advanced agents, deployment and failed correction to existential catastrophe.

  • First submission: June 16, 2022
  • Revision: August 13, 2024

Joseph Carlsmith examines one route to catastrophe: systems with advanced capabilities, planning and strategic awareness pursue problematic objectives. His six premises concern feasibility, development incentives, alignment difficulty, harmful deployment, permanent human disempowerment and the resulting loss of humanity’s long-term potential (§1).[1]

Power and correction

Power can help an agent achieve many objectives. Carlsmith distinguishes harmful mistakes from efforts to gain or retain control, and treats their connection as a substantive hypothesis (§§4.1–4.2). Correction could interrupt the pathway; §6.4 considers containment and institutional responses rather than assuming every failure becomes catastrophe.[1]

Estimates and limits

The original roughly 5% estimate by 2070 becomes above 10% in a May 2022 author note. These are subjective judgments, not measured frequencies. On p.47 he cautions that “the numbers here (and the exercise more broadly) should be held very lightly.”[1]

Historical context

The report identifies an April 2021 public release; arXiv submission followed in June 2022. The reviewed v2 is dated August 2024.[1]

As an editorial connection, Corrigibility addresses a link within this broader argument. A shutdown result alone leaves deployment incentives and institutional correction open. Current and Near-Term AI as a Potential Existential Risk Factor examines a different route through social systems.

Explore the chronology →

Sources

  1. Is Power-Seeking AI an Existential Risk? · Source record src-214 · Back to claim ↑1 ↑2 ↑3 ↑4

Last updated 2026-10-10