paper · Proceedings: 2024
Training Socially Aligned Language Models in Simulated Human Society
Uses simulated peer ratings, feedback and response revision as material for language-model alignment.
Can language models learn from iterative social feedback instead of isolated demonstrations? Ruibo Liu and coauthors introduce SANDBOX, a rule-guided society of language-model agents, and Stable Alignment.[1]
Method and contribution
Agents generate responses, gather peer ratings and feedback, and revise their answers. Stable Alignment uses these records in imitation, self-critique and realignment stages. The paper compares the trained models with baselines across alignment benchmarks and adversarial prompts, with ablations assessing the feedback and revision stages.[1]
Alignment relevance and limits
The approach makes simulated interaction a source of training feedback. It differs from Society-in-the-Loop: Programming the Algorithmic Social Contract, which concerns participation by affected people and institutions.
The authors note text-only interaction, primarily English evaluation and a static view of norms. Model-generated consensus can reproduce mainstream biases and omit underrepresented groups. Better benchmark ratings do not establish public legitimacy or represent real communities’ reflective endorsement.[1]
Historical context
Training Socially Aligned Language Models in Simulated Human Society identifies itself as an ICLR 2024 conference paper. The timeline uses year precision rather than assigning an unsupported release day.[1]
See Stable Alignment.
Sources
Pages that link here
- Key Papers: A Reading Guide note
- Ruibo Liu person
- Society-in-the-Loop concept
- Society-in-the-Loop: Programming the Algorithmic Social Contract paper
- Stable Alignment concept
Last updated 2026-10-08