person
Evan Hubinger
Coauthor of Risks from Learned Optimization in Advanced Machine Learning Systems.
Evan Hubinger is a coauthor of Risks from Learned Optimization in Advanced Machine Learning Systems. [1]
His contribution in that paper concerns Mesa-Optimization, Inner Alignment and the proposed Deceptive Alignment mechanism. This stub attributes a coauthored framework; it does not assign individual credit for particular sections.
Hubinger also coauthored Alignment faking in large language models. Its constructed experiments offer a distinct form of evidence from the conceptual analysis in his learned-optimization work.
Generalization and agent evaluation
Agentic Misalignment: How LLMs Could Be Insider Threats and Natural Emergent Misalignment from Reward Hacking in Production RL connect the earlier optimization taxonomy to controlled experiments on harmful agent behavior.
Sources
Pages that link here
Last updated 2026-10-08