concept

AI Safety via Debate

A proposed oversight method in which competing AI arguments help a human judge assess an answer.

Also known as: Debate, AI safety debate

Debate trains agents to argue competing answers before a human judge. The proposal aims to make difficult reasoning easier to check by focusing disagreement on smaller, assessable claims.[1]

It is a form of Scalable Oversight, not a guarantee that the most persuasive answer is true. The 2020 research writeup investigates protocols using human experts and non-expert judges.[2]

The 2020 work reports unreliable early judging and ambiguity failures; cross-examination was a proposed response, with unresolved training concerns.[2]

See the original proposal and the human protocol experiments.

Evidence and error localization

The original image experiment used machine search and a fixed classifier; its results do not measure human judging of unrestricted language.[1] The later human writeup asks whether debaters can actually locate an error small enough to expose. A capable judge still needs usable evidence from the debate.[2]

Sources

  1. AI safety via debate · Source record ref-debate-2018 · Back to claim ↑1 ↑2
  2. Writeup: Progress on AI Safety via Debate · Source record src-123 · Back to claim ↑1 ↑2 ↑3

Last updated 2026-10-09