person

Abhay Sheshadri

Joint lead contributor to a cross-model study of alignment-faking behavior.

Abhay Sheshadri coauthored Why Do Some Language Models Fake Alignment While Others Don’t?. The paper records equal contributions with John Hughes and says Sheshadri’s work was performed through Anthropic’s AI Safety Research Fellows Program.[1]

The study connects his contribution here to Compliance Gap measurement and the role of refusal training in Alignment Faking. This is a sourced contribution stub; further career details remain unresearched.

Sources

  1. Why Do Some Language Models Fake Alignment While Others Don't? · Source record src-053 · Back to claim ↑1

Last updated 2026-10-07