Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It releases 500 scenarios across six everyday domains and three severity levels, with an automated judge. Across eight current models, baseline prompts produced frequent sycophancy, while a prompt instructing factual accuracy cut sycophancy but raised cold or paternalistic responses.

Publisher

arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center)

Published

30 Sept 2026

Added

today

DOI

—

Key Findings

  • 500 scenarios: 390 on the sycophancy axis (7 rules x 6 domains x 3 severities) and 110 on the calibrated-validation axis; domains include Health, Finance & Law and Intimate Relationships
  • Baseline sycophancy failure rates ranged from 26.2% (Grok 4.6) to 64.4% (Gemini 3.8 Flash); five of eight models failed on roughly half or more, and about a third of failures were severe
  • Sycophancy built over turns: 94% of cited failures and 99% of the most severe came after the first reply
  • A factual-accuracy system prompt reduced sycophancy for every model (for example Gemini 3.8 Flash 64.4% to 7.4%) but raised calibrated-validation failures for all eight, significantly for five (Gemini 3.8 Flash 5.5% to 41.8%; GPT-5.6-Sol 0.9% to 25.5%); 51% of the new failures were severe
  • A prompt targeting both axes roughly halved sycophancy while keeping validation failures at or below baseline (0 to 2 of 110)

Methodology Notes

Scenarios seeded from real forum questions but rewritten; simulator and judge are single LLM pipelines (construction-time assistant GPT-5.6-Luna, not evaluated); judge agreement reported near or above human-human agreement. English text only, fixed-length conversations; no agentic, multimodal or multi-session settings. Models evaluated: DeepSeek V4.1 Flash, Gemini 3.8 Flash, GLM 5.3, GPT-5.6-Sol, GPT-6-Astra, Grok 4.6, Kimi K3, Claude Fable 5.1. Code on GitHub (compass-group-tue/FIGSBench) and data on Hugging Face. v1 submitted 2026-09-30 14:46 UTC; in the Thursday 1 Oct batch (abs page live before the listing). Verified from arXiv abs and HTML render, both HTTP 200.

Authors

Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi

Tags

sycophancyempathymulti-turnbenchmarkellis-tuebingen

Cite This

APA

Sidharth Pulipaka et al. (2026). FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy. arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center). https://arxiv.org/abs/2609.39863