Ask don't tell: Reducing sycophancy in large language models
Controlled experimental studies from the UK AI Security Institute isolating what provokes and prevents sycophancy in large language models. A nested factorial design compares questions to non-questions while varying epistemic certainty (statement, belief, conviction), perspective (I- vs user-perspective), and affirmation vs negation, measuring how sycophantically models phrase free-text responses. The framing findings are then used to build a mitigation: asking the model to convert non-questions into questions before answering.
Publisher
arXiv (UK AI Security Institute)
Published
27 Feb 2026
Added
yesterday
DOI
—
Key Findings
- Sycophancy is substantially higher in response to non-questions than to questions
- Sycophancy increases monotonically with the epistemic certainty conveyed by the user, and is amplified by I-perspective framing
- Asking a model to convert non-questions into questions before answering significantly reduces sycophancy, and the effect is stronger than a baseline prompt asking the model not to be sycophantic
- The framing effects generalise to choice sycophancy in a context-rich personalised setting: question vs statement framing shapes how often a model with extensive user knowledge picks the answer aligned with the user's stance in forced binary choices
Methodology Notes
Preprint, arXiv 2602.23971; v1 2026-02-27, revised through v4 2026-07-29. Affiliation UK AI Security Institute, London (read from the PDF title block; the abs page does not state affiliations). Controlled factorial experiments measuring expressed sycophancy in free-text responses; not yet peer-reviewed.
Sources
arXiv abstract page(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
Tags
Cite This
APA
Magda Dubois et al. (2026). Ask don't tell: Reducing sycophancy in large language models. arXiv (UK AI Security Institute). https://arxiv.org/abs/2602.23971
Related Insights
Measuring and Detecting Harmful AI Sycophancy
arXiv preprint · 6 Aug 2026
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026
When Do LLM Preferences Predict Downstream Behavior?
AI Security Institute (UK Department for Science, Innovation and Technology) · 26 Aug 2026
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic · 28 Aug 2026
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
arXiv (University of Illinois Chicago; National University of Singapore) · 27 Aug 2026