concept
AI Safety via Debate
A proposed oversight method in which competing AI arguments help a human judge assess an answer.
Also known as: Debate, AI safety debate
Debate trains agents to argue competing answers before a human judge. The proposal aims to make difficult reasoning easier to check by focusing disagreement on smaller, assessable claims.[1]
It is a form of Scalable Oversight, not a guarantee that the most persuasive answer is true. The 2020 research writeup investigates protocols using human experts and non-expert judges.[2]
The 2020 work reports unreliable early judging and ambiguity failures; cross-examination was a proposed response, with unresolved training concerns.[2]
See the original proposal and the human protocol experiments.
Evidence and error localization
The original image experiment used machine search and a fixed classifier; its results do not measure human judging of unrestricted language.[1] The later human writeup asks whether debaters can actually locate an error small enough to expose. A capable judge still needs usable evidence from the debate.[2]
Sources
- AI safety via debate · Source record ref-debate-2018 · Back to claim ↑1 ↑2
- Writeup: Progress on AI Safety via Debate · Source record src-123 · Back to claim ↑1 ↑2 ↑3
Pages that link here
Last updated 2026-10-09