concept

Stackelberg Game

A sequential game in which a leader chooses while anticipating a follower’s response.

A Stackelberg game models a leader choosing a strategy while anticipating a follower’s best response. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback uses this structure for coupled language-model and preference-model training.[1]

The equilibrium is a property of the specified game and utilities. It does not certify AI Alignment, the adequacy of the preference model, or convergence of a particular training procedure.

Equilibrium and training

A follower best response depends on the leader’s chosen strategy. In the Brown STA-RLHF formulation, a strong equilibrium resolves ties among follower best responses in the leader’s favor. The training algorithm approximates this sequential response with an inner preference-model update followed by a language-model update. A finite training run and a favorable synthetic reward are evidence about that procedure, not a convergence theorem.[1]

Sources

  1. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback · Source record src-041 · Back to claim ↑1 ↑2

Last updated 2026-10-09