Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-derived questions, totalling about 1.2 million trials. Finds conversational context, not model identity, is the dominant driver of sycophantic concession.
Publisher
arXiv (Virginia Tech)
Published
2 Aug 2026
Added
2 months ago
Key Findings
- Sycophancy rates vary about 67-fold across questions but only about 3-fold across models, indicating a single per-model sycophancy rate obscures the dominant sources of variance
- Fabricated authority sources roughly double concession when presented alongside the question but halve it when introduced after the model has already answered
- Chain-of-thought traces show models that re-examine their own prior answer tend to concede, while models that reason about the underlying medical facts hold their answers
Methodology Notes
Fully crossed factorial design: four conversational factors × five open-weight models × 500 MedQuAD-derived questions (~1.2M trials), with chain-of-thought analysis of concession reasoning. Preprint (arXiv 2608.01017, v1 2026-08-02), not yet peer-reviewed; title, authors, and date verified via the arXiv API.
Sources
arXiv abstract(opens in a new tab) (primary)
EMNLP 2026 programme (paper 5991-FIND, Findings of EMNLP 2026)(opens in a new tab) (29 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Kaike Ping, Buse Çarık, Caleb Wohn, Xiaohan Ding, Tongshuai Wang, Eugenia Rho
Tags
Cite This
APA
Kaike Ping et al. (2026). Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy. arXiv (Virginia Tech). https://arxiv.org/abs/2608.01017
Related Insights
SycEval: Evaluating LLM Sycophancy
arXiv (Stanford-led) · 12 Feb 2025
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California) · 15 Apr 2025
Affective Context Amplifies Sycophancy in LLM Responses
arXiv (preprint) · 21 Aug 2026
Measuring and Detecting Harmful AI Sycophancy
arXiv preprint · 6 Aug 2026
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics) · 3 Sept 2026
The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026) · 1 Mar 2026