Mitigating Social Sycophancy via Pluralistic Preference Optimization
Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer responses acceptable to all of them, without ground-truth labels. Evaluated on four social-sycophancy datasets covering personal advice, statements of intent to cause harm and r/AmITheAsshole verdicts.
Publisher
arXiv (Stanford University; University of Washington; Amazon)
Published
1 Oct 2026
Added
today
DOI
—
Key Findings
- On statements of intent to cause harm (PAS), PlurPO reduced the action endorsement rate by 89% on average across four base models
- On general advice questions (OEQ), the gap between model and human endorsement rates fell from 17.8% to 8.0% on average
- On AITA, PlurPO had the lowest false-negative (sycophantic) rate but a higher false-positive (over-critical) rate; it had the highest macro-F1
- Preference data built with Qwen3-8B transferred to Qwen3-32B; the procedure tuned for Qwen3-8B failed to train Llama 3.1 8B, whose simulated stakeholders vetoed every candidate response, so prompts and parameters were adapted per model family
Methodology Notes
Base models Qwen3-8B (primary), Phi-4, Llama 3.1 8B and Granite-4.1-8B; Iterative RPO training; held-out evaluation sets of 1,000 prompts per dataset; gpt-5-mini as LM judge using the human-validated endorsement prompt of Cheng et al.; six baselines including prompting and GEPA-optimised prompts. Datasets from prior social-sycophancy work (OEQ, PAS, AITA, AITA-Flipped). No human-participant outcome measured. v1 2026-10-01 23:00 UTC; not peer reviewed.
Sources
arXiv preprint(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta, Natasha Jaques, Emma Brunskill
Tags
Cite This
APA
Stephane Hatgis-Kessell et al. (2026). Mitigating Social Sycophancy via Pluralistic Preference Optimization. arXiv (Stanford University; University of Washington; Amazon). https://arxiv.org/abs/2610.02568