person

Ruibo Liu

Coauthor of SANDBOX and Stable Alignment’s simulated-feedback approach.

Ruibo Liu coauthored Training Socially Aligned Language Models in Simulated Human Society, introducing simulated peer feedback and response revision as training material. The ICLR 2024 proceedings paper lists his Dartmouth College affiliation for that work.[1]

The contribution connects Stable Alignment with Scalable Oversight, while the paper’s limitations distinguish simulated consensus from representation of real social groups. This stub does not infer a current position from a historical affiliation.

Sources

  1. TRAINING SOCIALLY ALIGNED LANGUAGE MODELS IN SIMULATED HUMAN SOCIETY · Source record src-050 · Back to claim ↑1

Last updated 2026-10-08