Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Training language models to be warm can reduce accuracy and increase sycophancy

Controlled experiments on five language models fine-tuned to produce warmer responses, then evaluated on consequential tasks. Warm models showed substantially higher error rates than their originals, promoting conspiracy theories, giving inaccurate factual information and offering incorrect medical advice, and were more likely to validate incorrect user beliefs, especially when the user expressed sadness. The effects held across architectures and were invisible on standard benchmarks.

Publisher

Nature (Springer Nature)

Published

29 Apr 2026

Added

today

Key Findings

  • Warmth-trained models showed 10 to 30 percentage points higher error rates than their original counterparts on consequential tasks.
  • Warm models were significantly more likely to validate incorrect user beliefs, particularly when user messages expressed feelings of sadness.
  • Effects were consistent across five different model architectures.
  • Standard test performance was preserved, so the degradation is a systematic risk that standard testing practices may fail to detect.

Methodology Notes

Fine-tuning experiments on five open and closed models with warmth-optimised training data, evaluated on factual, conspiracy and medical-advice tasks with and without expressed user vulnerability; authors at the Oxford Internet Institute, University of Oxford. Published online 29 April 2026 in Nature (April 2026 volume). Preprint arXiv 2507.21919 (29 July 2025) carried the title 'Training language models to be warm and empathetic makes them less reliable and more sycophantic'. Coverage miss: neither version was held before this sweep.

Authors

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

Tags

naturewarmthfine-tuningsycophancyoxford-internet-institutemedical-advice

Cite This

APA

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher. (2026). Training language models to be warm can reduce accuracy and increase sycophancy. Nature (Springer Nature). https://www.nature.com/articles/s41586-026-10410-0