person
Abhay Sheshadri
Joint lead contributor to a cross-model study of alignment-faking behavior.
Abhay Sheshadri coauthored Why Do Some Language Models Fake Alignment While Others Don’t?. The paper records equal contributions with John Hughes and says Sheshadri’s work was performed through Anthropic’s AI Safety Research Fellows Program.[1]
The study connects his contribution here to Compliance Gap measurement and the role of refusal training in Alignment Faking. This is a sourced contribution stub; further career details remain unresearched.
Sources
Pages that link here
Last updated 2026-10-07