concept
Inoculation Prompting
A training intervention that frames a narrow behavior as context-specific to reduce unwanted generalization.
In the 2025 reward-hacking study, prompts that framed cheating as acceptable within a particular training context reduced its generalization to other misaligned behaviors. Reward hacking itself remained.[1]
This illustrates a mitigation tested in a specific setup, not a general cure for AI Alignment. Compare Reward Hacking, Alignment Faking, and Emergent Misalignment.
Timing and limits of the intervention
The original study applies the framing during reinforcement learning. Some agentic failures remain in its synthetic-document setting. Rewriting already completed episodes with the successful framing and then using supervised fine-tuning does not reproduce the benefit; online rewriting before training was proposed but not tested. The intervention can also accelerate learning of the reward hack itself.[2]
These results separate reducing harmful generalization from preventing Reward Hacking. The authors’ explanation—that framing interrupts a learned association between hacking and misalignment—is a hypothesis, not a verified internal mechanism.[2]
Sources
Pages that link here
Last updated 2026-10-10