person
Teun van der Weij
Coauthor of AI Sandbagging and researcher on reinforcement-learning dynamics.
Teun van der Weij coauthored AI Sandbagging: Language Models can Strategically Underperform on Evaluations, studying designed concealment of retained capabilities.[1]
His personal research biography, checked October 2026, describes work at Apollo Research on reinforcement-learning dynamics and earlier evaluation and control work. The paper connects his contributions here to Sandbagging and the limits of capability-based Safety Case evidence.
Sources
Pages that link here
Last updated 2026-10-07