Skip to main content
Preprint Credible

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds a corpus of 290,460 labeled responses across 17 open- and closed-source LLMs in 12 everyday-advice domains, measures reversal rates, and tests whether harmful sycophancy can be detected automatically from response text alone.

Publisher

arXiv preprint

Published

6 Aug 2026

Added

1 week ago

DOI

Key Findings

  • Preference-induced stance reversal rates range from 5% to 56% across the 17 evaluated LLMs, with more capable models generally less sycophantic
  • Detecting harmful sycophancy from response text alone is feasible with trained detectors
  • Detection performance drops substantially on models unseen during detector training; preliminary mitigation strategies are proposed

Methodology Notes

arXiv:2608.05624 (cs.AI, cs.CL), v1 submitted 2026-08-06, marked under review. Author affiliations not stated on the arXiv listing; the author group is associated with Arizona State University's data-mining orbit (Huan Liu). 290,460 labeled responses, 17 LLMs, 12 advice domains.

Sources

arXiv abstract page (primary)

Archived snapshot (Wayback Machine) — preserved against link rot

Authors

Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu

Tags

sycophancy-detectionstance-reversalcorpusasu

Cite This

APA

Bohan Jiang et al. (2026). Measuring and Detecting Harmful AI Sycophancy. arXiv preprint. https://arxiv.org/abs/2608.05624