concept

Preference Aggregation

Combining multiple people’s preferences into a collective decision or objective.

Preference aggregation combines individual preferences into a decision rule. Dynamic value alignment through preference aggregation of multiple objectives compares majority and proportional rules in a limited traffic-control simulation.[2]

Choosing the rule is itself a question of Social Alignment. Aggregation does not automatically make a result fair, legitimate or resistant to strategic reporting. The Challenge of Value Alignment: from Fairer Algorithms to AI Safety frames the underlying problem of plural social values.[3]

What is being aggregated?

In the traffic study, users vote for system-wide objectives; trained networks supply action scores and the voting rule determines their weights. The composition of voters changes, while each simulated person’s preference remains fixed. An added objective requires a corresponding trained model, and the experiment tests only two objectives under a half-and-half population split.[2]

This differs from updating a model of a designer’s preferences. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback refreshes its preference model on the language model’s current outputs, but leaves preference aggregation as future work. Better feedback adaptation does not by itself decide how conflicting people’s preferences should count.[4]

Reference groups and disagreement

InstructGPT identifies labelers, researchers and customers as distinct influences on its target behavior. Most comparison tasks are labeled by one contractor; the paper warns that averaging preferences may fail to represent minority groups. Better prediction of rankings does not settle the choice of reference group or how disagreement should be resolved.[1]

Sources

  1. Training language models to follow instructions with human feedback · Source record src-179 · Back to claim ↑1
  2. Dynamic value alignment through preference aggregation of multiple objectives · Source record src-054 · Back to claim ↑1 ↑2
  3. The Challenge of Value Alignment: from Fairer Algorithms to AI Safety · Source record src-045 · Back to claim ↑1
  4. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback · Source record src-041 · Back to claim ↑1

Last updated 2026-10-09