concept

Goodhart's Law

Optimizing a proxy can weaken its relationship to the outcome it was meant to measure.

Goodhart’s law names a family of problems that arise when optimization pressures a proxy beyond the conditions where it usefully tracks a goal.

Manheim and Garrabrant distinguish several mechanisms rather than treating every metric failure as the same phenomenon. In AI Alignment, Reward Hacking is one way a successful score can accompany an unsuccessful task.

The institutional analogy is useful: a target can redirect behavior toward the measurement. It does not show that all metrics fail, or that every feedback loop must collapse.

Four ways a proxy can fail

Categorizing Variants of Goodhart’s Law separates four mechanisms. The following teaching examples illustrate its distinctions; they are not measured incidents.[1]

Mechanism What changes Illustrative example
Regressional Selecting a high noisy score also selects favorable measurement error. The highest-scoring candidate on a noisy test may owe part of that lead to luck.
Extremal Optimization reaches a region where the fitted relationship no longer holds. A reward model fitted to ordinary answers is used to rank increasingly unusual outputs.
Causal An intervention changes the process connecting the metric to the goal. A teacher directly raises recorded grades without improving learning.
Adversarial Another actor responds to the metric with a different objective. A rewarded participant finds a way to raise the evaluator’s score while missing its purpose.

In the regressional model, selecting a high score can still improve the true goal: the goal improves less than the score suggests. Extremal failures include both an inadequate model and a change of regime. The paper’s causal category requires an intervention; simply overlooking a common cause while selecting candidates can instead produce regressional or extremal failure. Adversarial responses can exploit any of these mechanisms, so the categories can overlap.[1]

Manheim and Garrabrant caution: “These varied forms often occur together, but defining them individually is useful.” (v4, p. 2.)[1]

Diagnose before choosing a defense

Ask what the proxy measures, where its relationship was tested, what actions can change the measurement process, and who responds to the incentive. This is an editorial checklist derived from the taxonomy, rather than a validated mitigation procedure. Better measurement addresses a different weakness from preventing direct score alteration; neither guarantees success across the other mechanisms.[1]

Reward Hacking provides concrete research examples of proxy optimization. The taxonomy supplies possible explanations, not proof that every reward-hacking result has one uniquely identified cause. A proxy’s failure also does not by itself establish Deceptive Alignment or persistent intent.

Sources

  1. Categorizing Variants of Goodhart's Law · Source record ref-e178cbf073c0 · Back to claim ↑1 ↑2 ↑3 ↑4

Last updated 2026-10-10