FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It releases 500 scenarios across six everyday domains and three severity levels, with an automated judge. Across eight current models, baseline prompts produced frequent sycophancy, while a prompt instructing factual accuracy cut sycophancy but raised cold or paternalistic responses.
Publisher
arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center)
Published
30 Sept 2026
Added
today
DOI
—
Key Findings
- 500 scenarios: 390 on the sycophancy axis (7 rules x 6 domains x 3 severities) and 110 on the calibrated-validation axis; domains include Health, Finance & Law and Intimate Relationships
- Baseline sycophancy failure rates ranged from 26.2% (Grok 4.6) to 64.4% (Gemini 3.8 Flash); five of eight models failed on roughly half or more, and about a third of failures were severe
- Sycophancy built over turns: 94% of cited failures and 99% of the most severe came after the first reply
- A factual-accuracy system prompt reduced sycophancy for every model (for example Gemini 3.8 Flash 64.4% to 7.4%) but raised calibrated-validation failures for all eight, significantly for five (Gemini 3.8 Flash 5.5% to 41.8%; GPT-5.6-Sol 0.9% to 25.5%); 51% of the new failures were severe
- A prompt targeting both axes roughly halved sycophancy while keeping validation failures at or below baseline (0 to 2 of 110)
Methodology Notes
Scenarios seeded from real forum questions but rewritten; simulator and judge are single LLM pipelines (construction-time assistant GPT-5.6-Luna, not evaluated); judge agreement reported near or above human-human agreement. English text only, fixed-length conversations; no agentic, multimodal or multi-session settings. Models evaluated: DeepSeek V4.1 Flash, Gemini 3.8 Flash, GLM 5.3, GPT-5.6-Sol, GPT-6-Astra, Grok 4.6, Kimi K3, Claude Fable 5.1. Code on GitHub (compass-group-tue/FIGSBench) and data on Hugging Face. v1 submitted 2026-09-30 14:46 UTC; in the Thursday 1 Oct batch (abs page live before the listing). Verified from arXiv abs and HTML render, both HTTP 200.
Sources
arXiv preprint(opens in a new tab) (primary)
FIGSBench dataset (Hugging Face)(opens in a new tab) (30 Sept 2026)
FIGSBench code (GitHub)(opens in a new tab) (30 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi
Tags
Cite This
APA
Sidharth Pulipaka et al. (2026). FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy. arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center). https://arxiv.org/abs/2609.39863
Related Insights
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
arXiv (University of Illinois Chicago; National University of Singapore) · 27 Aug 2026
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
arXiv (Texas A&M University; University of Cincinnati) · 8 Sept 2026
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025
Training language models to be warm can reduce accuracy and increase sycophancy
Nature (Springer Nature) · 29 Apr 2026