paper · First submission: May 2, 2018
AI safety via debate
Proposes competing AI arguments as a way for people to evaluate answers they could not produce or check unaided.
- Revision: October 22, 2018
Can a person judge an answer that they cannot independently produce? Geoffrey Irving, Paul Christiano and Dario Amodei propose training competing agents to debate before a human judge.[1]
Contribution and method
Each debater has an incentive to expose weaknesses in the other’s answer. The paper motivates this Scalable Oversight protocol with an idealized computational argument and an initial image-classification experiment in which agents reveal evidence to a constrained judge.[1]
What the initial experiment tested
In section 3.1 of the reviewed arXiv v2, a fixed classifier judges MNIST digits from four or six revealed nonzero pixels. One player is assigned the true label, the other an incorrect label; pixels themselves cannot be fabricated. Debaters use Monte Carlo Tree Search with access to the judge, rather than learned natural-language policies.[1]
Table 2 distinguishes precommitting to an incorrect label from adapting it during play. With six pixels, the honest win rate averages 88.9% with precommitment and 74.4% without it; random-pixel judge accuracy is 59.4%. These are restricted-task results. Section 3.2 describes informal human cat-versus-dog play and leaves formal experiments to future work.[1]
Alignment relevance and limits
The associated AI Safety via Debate concept concerns making difficult claims checkable through adversarial discussion. It addresses the costly-supervision problem in Concrete Problems in AI Safety.
The theoretical result depends on idealized play and judge behavior. A restricted image task does not show that real people can reliably adjudicate arbitrary scientific or moral questions, or that trained agents will reach the intended equilibrium. Persuasion, shared errors and judge limitations remain concerns.[1]
See AI safety via debate and later experimental work.
Historical context
The first arXiv version of AI safety via debate was submitted on May 2, 2018. The paper proposes training competing agents to debate before a human judge.[1]
The reviewed version is the October 22, 2018 revision; the chronology retains first submission. No complete v1/v2 comparison was performed.
Sources
- AI safety via debate · Source record ref-debate-2018 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6
Pages that link here
- AI Safety via Debate concept
- Dario Amodei person
- Key Papers: A Reading Guide note
- Paul Christiano person
- Scalable Oversight concept
- Writeup: Progress on AI Safety via Debate paper
Last updated 2026-10-09