The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedical questions (PubMedQA). Introduces Resistance, a turn-level survival-style metric for the share of cases that have not yet flipped at each pushback, and a Sticky Incorrect Ratio comparing how often models preserve correct versus incorrect answers. Gemini models resist best, OpenAI reasoning models next, and Claude models flip under even mild pushback; models flip more readily on clear-answer questions than on ambiguous diagnoses and abandon correct answers more often than incorrect ones.
Publisher
Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026)
Published
1 Mar 2026
Added
today
Key Findings
- On MedCaseReasoning, Gemini models kept Resistance above 80% after four pushbacks; Claude Haiku 4.5 ended at 4.7%, meaning more than 95% of its diagnoses had flipped, and Claude models' largest drop came at the first, mildest pushback
- On PubMedQA every model did worse: Gemini models ended at 35.4% to 49.7% Resistance while OpenAI reasoning models (o3-mini, o4-mini, GPT-5) all finished below 20%
- Models abandon correct answers more readily than incorrect ones: by turn four OpenAI models flipped 88-98% of initially correct answers but 68-87% of initially incorrect ones; GPT-5's gap was 20 percentage points and GPT-4.1 retained only 2.5% of its correct answers
- Claude Sonnet 4.5 was the only model more likely to abandon incorrect than correct answers (Sticky Incorrect Ratio 0.7), though with high overall flip rates
- Pushbacks escalate from mild confusion through reassertion and anecdotal counterexample to direct challenge, generated by o4-mini; Gemini and OpenAI reasoning models fell most at the anecdotal counterexample, Claude and OpenAI non-reasoning models at the first pushback
- Response patterns ('Yes, but...' versus 'Yes, and...') predicted flips better than specific phrases
Methodology Notes
ACL Anthology 2026.healing-1.2, DOI 10.18653/v1/2026.healing-1.2, pages 19-34, HeaLing 2026 workshop at LREC (Rabat, March 2026; month precision). Affiliations from the PDF title block: Stanford University, Harvard Medical School and Scripps Research. Models: Gemini 3 Pro, Gemini 2.5 Pro and Flash; GPT-5, GPT-4.1, GPT-4o, o4-mini, o3-mini; Claude models including Haiku 4.5 and Sonnet 4.5. Pushbacks and flip judgments are generated and scored by o4-mini (Gemini 2.5 Pro for its own family), so the pressure is simulated rather than from real patients; the setting is a medical assistant answering clinical questions, closer to clinician-facing QA than to a patient self-managing care. Claude models initially refused PubMedQA prompts as medical-advice requests and the system prompts were modified. No data or code link found in the PDF.
Sources
ACL Anthology(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Taeil Matthew Kim, Luyang Luo, Sung Eun Kim, Arjun Kumar Manrai, Eric Topol, Pranav Rajpurkar
Tags
Cite This
APA
Taeil Matthew Kim et al. (2026). The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations. Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026). https://aclanthology.org/2026.healing-1.2/
Related Insights
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
arXiv (Virginia Tech) · 2 Aug 2026
SycEval: Evaluating LLM Sycophancy
arXiv (Stanford-led) · 12 Feb 2025
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
arXiv (University of Illinois Chicago; National University of Singapore) · 27 Aug 2026
Ask don't tell: Reducing sycophancy in large language models
arXiv (UK AI Security Institute) · 27 Feb 2026