person

Teun van der Weij

Coauthor of AI Sandbagging and researcher on reinforcement-learning dynamics.

Teun van der Weij coauthored AI Sandbagging: Language Models can Strategically Underperform on Evaluations, studying designed concealment of retained capabilities.[1]

His personal research biography, checked October 2026, describes work at Apollo Research on reinforcement-learning dynamics and earlier evaluation and control work. The paper connects his contributions here to Sandbagging and the limits of capability-based Safety Case evidence.

Sources

  1. AI Sandbagging: Language Models can Strategically Underperform on Evaluations · Source record src-011 · Back to claim ↑1
  2. Teun van der Weij: personal research biography · Source record ref-4d7d4e1d7f97

Last updated 2026-10-07