note
Who Supplies Alignment Feedback?
Comparing human judges, simulated social feedback, and negotiation among affected people.
These approaches address different parts of alignment. Comparing their feedback sources helps explain what each can contribute.
| Approach | Who supplies feedback? | Contribution and limit |
|---|---|---|
| AI Safety via Debate | A human judges competing AI arguments. | A proposal for evaluating difficult work; its success depends on the debate mechanism and judge. [1] |
| Stable Alignment | Simulated agents provide ratings and responses are revised. | A training method with benchmark results; simulated feedback does not establish public agreement. [2] |
| Society-in-the-Loop | Affected stakeholders negotiate values and monitor outcomes. | A conceptual governance framework; it does not supply a ready-made training objective. [3] |
Scalable Oversight asks how to evaluate difficult work. Social Alignment also asks whose interests that evaluation represents. The reading guide connects both questions to the wider literature.
Sources
- AI safety via debate · Source record ref-debate-2018 · Back to claim ↑1
- TRAINING SOCIALLY ALIGNED LANGUAGE MODELS IN SIMULATED HUMAN SOCIETY · Source record src-050 · Back to claim ↑1
- Society-in-the-Loop: Programming the Algorithmic Social Contract · Source record ref-ad43ed014e22 · Back to claim ↑1
Pages that link here
Last updated 2026-10-07