paper · First submission: October 20, 2023
Towards Understanding Sycophancy in Language Models
Controlled tests connect agreeable but inaccurate answers with imperfect human preference signals.
- Proceedings: 2024
- Revision: May 10, 2025
Mrinank Sharma, Meg Tong and colleagues examine Sycophancy in five assistants: Claude 1.3, Claude 2, GPT-3.5, GPT-4 and LLaMA 2. Tests change user cues while probing feedback, answer revision, biased answers and repetition of user mistakes (§3).[1]
What the evidence suggests
Preference-data analysis and preference-model experiments suggest approval signals can reward agreement over accuracy. Optimization has mixed effects across tasks and methods; the study does not show that every use of human feedback increases sycophancy (§4).[1]
Why read it here?
It helps explain why human feedback requires scrutiny. The 2025 rollback offers a later deployment comparison, not replication of these experiments.
Historical context
First arXiv submission was October 20, 2023. The reviewed v4 PDF labels ICLR 2024 publication and a May 10, 2025 revision. Selected main sections were inspected; code, full appendices and version equivalence were not audited. These controlled tasks do not measure everyday incident prevalence.[1]
Sources
Pages that link here
Last updated 2026-10-11