paper · Research release: February 5, 2020
Writeup: Progress on AI Safety via Debate
Reports human debate experiments, ambiguity failures and proposed cross-examination protocols for non-expert oversight.
Beth Barnes and Paul Christiano report OpenAI’s Reflection-Humans work from late 2019, developing AI Safety via Debate for non-expert oversight.[1]
Human experiments
Using tricky physics questions, expert debaters defended a correct answer and an inferior alternative. Early debates remained unreliable even with motivated judges spending about an hour. The later target—over 90% correct judgments within ten minutes—was an objective, not an achieved result.[1]
Protocol changes
Dishonest debaters could evade questions or exploit ambiguous claims. Recursive debates narrowed disagreement; proposed cross-examination queried noncommunicating copies of an opponent to expose inconsistent interpretations without consuming judge attention. Human approximations used backtracking or paired teams; the latter was untested.[1]
Why finding a flaw matters
The writeup’s “Long computation problem” distinguishes knowing an answer is wrong from finding a small error a judge can assess. If neither debater understands a long supporting argument well enough to isolate its flaw, a dishonest debater may force a draw. Recursion alone does not supply the missing ability.[1]
The “Current concerns” section also questions how judges should update after objections are conceded. Debaters choose objections in light of their winning chances, so concessions do not straightforwardly certify truth. The authors worry that rewarding completely honest play might fail to provide a useful training gradient toward less dishonesty. These are unresolved protocol and learning concerns, not demonstrated failures of trained superhuman models.[1]
Significance and limits
The writeup develops the 2018 proposal toward practical Scalable Oversight. Unresolved issues include judging conceded objections, training dynamics and arguments too complex for either debater to locate an error. Human experiments do not establish reliable superhuman oversight.[1]
Historical context
Released February 5, 2020; the research occurred in Q3–Q4 2019.[1]
Sources
- Writeup: Progress on AI Safety via Debate · Source record src-123 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7
Pages that link here
- AI safety via debate paper
- AI Safety via Debate concept
- Scalable Oversight concept
Last updated 2026-10-09