concept
Stable Alignment
A method that trains language models using feedback and revised responses from simulated social interactions.
Liu and coauthors use a simulated society of language-model agents to generate ratings, feedback, and revised responses for training. They call their framework Stable Alignment.[1]
The paper reports improvements on its alignment evaluations. Simulated agreement does not by itself establish that real communities endorse the resulting norms. See Social Alignment, Training Socially Aligned Language Models in Simulated Human Society, and Training Socially Aligned Language Models in Simulated Human Society.
Ruibo Liu and coauthors explicitly limit their reported results to primarily English, text-based settings and warn that simulated feedback can miss underrepresented communities. Preference Aggregation addresses a related but different problem: combining preferences in a decision rule.
Sources
Pages that link here
Last updated 2026-10-08