concept
Stackelberg Game
A sequential game in which a leader chooses while anticipating a follower’s response.
A Stackelberg game models a leader choosing a strategy while anticipating a follower’s best response. STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback uses this structure for coupled language-model and preference-model training.[1]
The equilibrium is a property of the specified game and utilities. It does not certify AI Alignment, the adequacy of the preference model, or convergence of a particular training procedure.
Equilibrium and training
A follower best response depends on the leader’s chosen strategy. In the Brown STA-RLHF formulation, a strong equilibrium resolves ties among follower best responses in the leader’s favor. The training algorithm approximates this sequential response with an inner preference-model update followed by a language-model update. A finite training run and a favorable synthetic reward are evidence about that procedure, not a convergence theorem.[1]
Sources
Pages that link here
Last updated 2026-10-09