Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unethical-request items, with a judge scoring position strength each turn. Evaluates four production systems and three Olmo-3-7B variants and reports collapse rates as a function of conversation length, tactic and user-proxy adaptivity.

Publisher

arXiv (Texas A&M University; University of Cincinnati)

Published

8 Sept 2026

Added

today

Key Findings

  • Collapse rates increase with conversation length for every model tested; short-horizon protocols underestimate sycophancy and resistance under sustained pressure remains unreliable across current models.
  • For models with accessible reasoning traces, the correct position often remains represented in the trace while the visible reply concedes, which the authors read as the model choosing to please the user rather than lacking the knowledge.
  • An adaptive LLM user proxy exposes more sycophantic collapse than pre-generated scripted pushback.
  • Among tactics, emotional appeals are most associated with inducing collapse.
  • Production systems tested are Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.1 Pro and DeepSeek V4 Pro; Claude Sonnet 5 serves as the judge, with 86% human-judge verdict agreement on a 50-item stratified check; code and data are released.

Methodology Notes

200 items (100 false presuppositions, 100 unethical queries); target models: four production systems plus Olmo-3-7B Base, Instruct and Think; judge is Claude Sonnet 5 with a human check on 50 verdicts (86% agreement); no real users; the judge and one target share a developer. arXiv 2609.09090 version 1, 8 September 2026 (announced 9 September); not peer reviewed.

Authors

Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang

Tags

spinemulti-turnadaptive-user-proxysycophancytexas-a-m

Cite This

APA

Leyuan Tang et al. (2026). Measuring LLM Sycophancy under Sustained Multi-Turn Pressure. arXiv (Texas A&M University; University of Cincinnati). https://arxiv.org/abs/2609.09090