person

Monte MacDiarmid

Monte MacDiarmid is a core research contributor to the reward-hacking generalization study.

Monte MacDiarmid is a core research contributor to the reward-hacking generalization study, documented by the original paper.[1]

Contribution

Natural Emergent Misalignment from Reward Hacking in Production RL describes the contribution and its limitations. This is a contribution-focused stub; current affiliations have not been independently established.

Sources

  1. Natural emergent misalignment from reward hacking in production RL · Source record src-037 · Back to claim ↑1

Last updated 2026-10-07