Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Reddit moral-advice data behind the ELEPHANT evaluation, the authors show that responses scored as more socially sycophantic are also more receptive, and that rewriting human answers to be more receptive while keeping their conclusions raises their sycophancy scores. A preregistered experiment with 200 participants finds people prefer the more receptive versions. A prompting approach that first anchors the model to its own judgment raises receptiveness without raising substantive deference.
Publisher
arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University)
Published
22 Sept 2026
Added
today
DOI
—
Key Findings
- Responses classified as more socially sycophantic on the ELEPHANT moral-advice data were also more receptive; the two measures were tightly coupled
- Rewriting human-written responses to be more receptive while preserving the substantive conclusion caused them to be scored as more socially sycophantic
- In a preregistered experiment (n=200), participants preferred the more receptive of two substantively equivalent responses, expected users to be more likely to listen to them, and were more willing to seek advice from their authors; the pattern held among participants who judged the original asker to be in the wrong
- Model estimates were computed over 949 (GPT-5.6 Terra), 869 (Claude Sonnet 5), 975 (Gemini 3.7 Flash) and 350 (Llama 4 Scout) examples spanning 1,339 unique posts
- Anchoring a model to its independent judgment before asking for a receptive answer increased receptiveness without increasing substantive deference; prompting for receptiveness alone also increased deference
Methodology Notes
Observational analysis on the AITA-YTA subset of the ELEPHANT dataset (2,000 posts, 1,892 after filtering) with automated receptiveness and social-sycophancy scoring; controlled rewrites; preregistered online survey experiment (n=200) measuring stated preferences; four models (GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.7 Flash, Llama 4 Scout). Author-stated limits: moral-advice domain only; automated receptiveness proxy; rewrites may differ in fluency; stated rather than behavioural outcomes; deference measured against one user signal. v1 posted 2026-09-22; code and data on GitHub.
Sources
Authors
Calvin Isley, Johann D. Gaebler, Max Lamparth, Julia Minson, Sharad Goel
Tags
Cite This
APA
Calvin Isley et al. (2026). Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models. arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University). https://arxiv.org/abs/2609.26579
Related Insights
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025
Towards Understanding Sycophancy in Language Models
Anthropic · 20 Oct 2023
Feeling Right vs. Being Right: How AI Sycophancy Affects Value-Laden Deliberation
Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Seoul National University; Taejae University; Yonsei University · 1 Jul 2026