person
Ruibo Liu
Coauthor of SANDBOX and Stable Alignment’s simulated-feedback approach.
Ruibo Liu coauthored Training Socially Aligned Language Models in Simulated Human Society, introducing simulated peer feedback and response revision as training material. The ICLR 2024 proceedings paper lists his Dartmouth College affiliation for that work.[1]
The contribution connects Stable Alignment with Scalable Oversight, while the paper’s limitations distinguish simulated consensus from representation of real social groups. This stub does not infer a current position from a historical affiliation.
Sources
Pages that link here
Last updated 2026-10-08