paper · Proceedings: 2024

Training Socially Aligned Language Models in Simulated Human Society

Uses simulated peer ratings, feedback and response revision as material for language-model alignment.

Can language models learn from iterative social feedback instead of isolated demonstrations? Ruibo Liu and coauthors introduce SANDBOX, a rule-guided society of language-model agents, and Stable Alignment.[1]

Method and contribution

Agents generate responses, gather peer ratings and feedback, and revise their answers. Stable Alignment uses these records in imitation, self-critique and realignment stages. The paper compares the trained models with baselines across alignment benchmarks and adversarial prompts, with ablations assessing the feedback and revision stages.[1]

Alignment relevance and limits

The approach makes simulated interaction a source of training feedback. It differs from Society-in-the-Loop: Programming the Algorithmic Social Contract, which concerns participation by affected people and institutions.

The authors note text-only interaction, primarily English evaluation and a static view of norms. Model-generated consensus can reproduce mainstream biases and omit underrepresented groups. Better benchmark ratings do not establish public legitimacy or represent real communities’ reflective endorsement.[1]

Historical context

Training Socially Aligned Language Models in Simulated Human Society identifies itself as an ICLR 2024 conference paper. The timeline uses year precision rather than assigning an unsupported release day.[1]

See Stable Alignment.

Explore the chronology →

Sources

  1. TRAINING SOCIALLY ALIGNED LANGUAGE MODELS IN SIMULATED HUMAN SOCIETY · Source record src-050 · Back to claim ↑1 ↑2 ↑3 ↑4

Last updated 2026-10-08