Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

SycEval: Evaluating LLM Sycophancy

A framework for quantifying progressive and regressive sycophancy in LLMs (GPT-4o, Claude-Sonnet, Gemini-1.5-Pro) across math (AMPS) and medical (MedQuad) tasks under user rebuttal pressure. It measures how often models change correct answers when challenged.

Publisher

arXiv (Stanford-led)

Published

12 Feb 2025

Added

3 months ago

Key Findings

  • Sycophancy occurred in 58.2% of cases, with 78.5% persistence across turns
  • Preemptive rebuttals triggered more sycophancy than in-context rebuttals
  • Distinguishes progressive (toward correct) from regressive (away from correct) sycophancy

Methodology Notes

Preprint (arXiv 2502.08177, 2025-02-12). Rebuttal-pressure protocol over math and medical QA; measures answer stability rather than emotional/relational sycophancy.

Authors

Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin, Anson Zhou, Roxana Daneshjou, Sanmi Koyejo

Tags

arxivsycevalsycophancybenchmarkstanford

Cite This

APA

Aaron Fanous et al. (2025). SycEval: Evaluating LLM Sycophancy. arXiv (Stanford-led). https://arxiv.org/abs/2502.08177