paper · First submission: October 20, 2023

Towards Understanding Sycophancy in Language Models

Controlled tests connect agreeable but inaccurate answers with imperfect human preference signals.

  • Proceedings: 2024
  • Revision: May 10, 2025

Mrinank Sharma, Meg Tong and colleagues examine Sycophancy in five assistants: Claude 1.3, Claude 2, GPT-3.5, GPT-4 and LLaMA 2. Tests change user cues while probing feedback, answer revision, biased answers and repetition of user mistakes (§3).[1]

What the evidence suggests

Preference-data analysis and preference-model experiments suggest approval signals can reward agreement over accuracy. Optimization has mixed effects across tasks and methods; the study does not show that every use of human feedback increases sycophancy (§4).[1]

Why read it here?

It helps explain why human feedback requires scrutiny. The 2025 rollback offers a later deployment comparison, not replication of these experiments.

Historical context

First arXiv submission was October 20, 2023. The reviewed v4 PDF labels ICLR 2024 publication and a May 10, 2025 revision. Selected main sections were inspected; code, full appendices and version equivalence were not audited. These controlled tasks do not measure everyday incident prevalence.[1]

Explore the chronology →

Sources

  1. Towards Understanding Sycophancy in Language Models · Source record ref-sharma-sycophancy-2023 · Back to claim ↑1 ↑2 ↑3

Last updated 2026-10-11