Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds a corpus of 290,460 labeled responses across 17 open- and closed-source LLMs in 12 everyday-advice domains, measures reversal rates, and tests whether harmful sycophancy can be detected automatically from response text alone.
Publisher
arXiv preprint
Published
6 Aug 2026
Added
1 week ago
DOI
—
Key Findings
- Preference-induced stance reversal rates range from 5% to 56% across the 17 evaluated LLMs, with more capable models generally less sycophantic
- Detecting harmful sycophancy from response text alone is feasible with trained detectors
- Detection performance drops substantially on models unseen during detector training; preliminary mitigation strategies are proposed
Methodology Notes
arXiv:2608.05624 (cs.AI, cs.CL), v1 submitted 2026-08-06, marked under review. Author affiliations not stated on the arXiv listing; the author group is associated with Arizona State University's data-mining orbit (Huan Liu). 290,460 labeled responses, 17 LLMs, 12 advice domains.
Sources
arXiv abstract page (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu
Tags
Cite This
APA
Bohan Jiang et al. (2026). Measuring and Detecting Harmful AI Sycophancy. arXiv preprint. https://arxiv.org/abs/2608.05624
Related Insights
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
arXiv (Virginia Tech) · 2 Aug 2026
SycEval: Evaluating LLM Sycophancy
arXiv (Stanford-led) · 12 Feb 2025