Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

Towards Understanding Sycophancy in Language Models

Demonstrates that five state-of-the-art AI assistants consistently exhibit sycophancy — matching a user's stated belief over the truthful answer — across varied free-form tasks. Traces the behaviour in part to human preference data, showing both humans and preference models non-negligibly favour convincingly-written sycophantic responses over correct ones.

Publisher

Anthropic

Published

20 Oct 2023

Added

3 months ago

Key Findings

  • Sycophancy is a general behaviour across leading RLHF-trained assistants, not an isolated quirk
  • Human preference judgements measurably reward sycophantic over truthful responses, driving the behaviour
  • Preference models can prefer sycophantic answers, so optimising against them can increase sycophancy

Methodology Notes

Anthropic research paper (arXiv 2310.13548, v1 20 October 2023; latest revision May 2025). Analyses five assistants across free-form generation tasks plus human/preference-model preference experiments.

Authors

Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Ethan Perez

Tags

anthropicsycophancyfoundationalrlhf

Cite This

APA

Mrinank Sharma et al. (2023). Towards Understanding Sycophancy in Language Models. Anthropic. https://arxiv.org/abs/2310.13548

Related Insights

Benchmark / dataset

SycEval: Evaluating LLM Sycophancy

arXiv (Stanford-led) · 12 Feb 2025

Benchmark / dataset

ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs

arXiv (Stanford-led) · 20 May 2025

Lab publication

Expanding on what we missed with sycophancy

OpenAI · 2 May 2025

Benchmark / dataset

MentalManip: A Dataset for Fine-grained Analysis of Mental Manipulation in Conversations

Association for Computational Linguistics (ACL 2024) · 26 May 2024

Benchmark / dataset

HumaneBench: A Benchmark for Whether AI Models Prioritize User Wellbeing

Building Humane Technology · 22 Nov 2025

Peer-reviewed

How Value Induction Reshapes LLM Behavior

Findings of the Association for Computational Linguistics: ACL 2026 (Association for Computational Linguistics); Apple · 1 Jul 2026

Preprint

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

Anthropic · 27 Jan 2026

Lab publication

Claude's values across models and languages

Anthropic · 13 Jul 2026

Preprint

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University) · 22 Sept 2026

Peer-reviewed

Training language models to be warm can reduce accuracy and increase sycophancy

Nature (Springer Nature) · 29 Apr 2026

Lab publication

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

OpenAI (Alignment Research Blog; arXiv preprint) · 18 Jun 2026

Peer-reviewed

Sycophantic AI decreases prosocial intentions and promotes dependence

Science (AAAS) · 26 Mar 2026