Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Mitigating Social Sycophancy via Pluralistic Preference Optimization

Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer responses acceptable to all of them, without ground-truth labels. Evaluated on four social-sycophancy datasets covering personal advice, statements of intent to cause harm and r/AmITheAsshole verdicts.

Publisher

arXiv (Stanford University; University of Washington; Amazon)

Published

1 Oct 2026

Added

today

DOI

—

Key Findings

  • On statements of intent to cause harm (PAS), PlurPO reduced the action endorsement rate by 89% on average across four base models
  • On general advice questions (OEQ), the gap between model and human endorsement rates fell from 17.8% to 8.0% on average
  • On AITA, PlurPO had the lowest false-negative (sycophantic) rate but a higher false-positive (over-critical) rate; it had the highest macro-F1
  • Preference data built with Qwen3-8B transferred to Qwen3-32B; the procedure tuned for Qwen3-8B failed to train Llama 3.1 8B, whose simulated stakeholders vetoed every candidate response, so prompts and parameters were adapted per model family

Methodology Notes

Base models Qwen3-8B (primary), Phi-4, Llama 3.1 8B and Granite-4.1-8B; Iterative RPO training; held-out evaluation sets of 1,000 prompts per dataset; gpt-5-mini as LM judge using the human-validated endorsement prompt of Cheng et al.; six baselines including prompting and GEPA-optimised prompts. Datasets from prior social-sycophancy work (OEQ, PAS, AITA, AITA-Flipped). No human-participant outcome measured. v1 2026-10-01 23:00 UTC; not peer reviewed.

Authors

Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta, Natasha Jaques, Emma Brunskill

Tags

social-sycophancyrelationship-adviceaitapreference-optimizationstanford

Cite This

APA

Stephane Hatgis-Kessell et al. (2026). Mitigating Social Sycophancy via Pluralistic Preference Optimization. arXiv (Stanford University; University of Washington; Amazon). https://arxiv.org/abs/2610.02568