When Do LLM Preferences Predict Downstream Behavior?
Study by the UK AI Security Institute testing whether large language models' elicited preferences predict their downstream behavior. Five frontier LLMs were evaluated across three domains: donation advice given to simulated users, refusal behavior, and task performance. The models acted on their own preferences in advice and refusal settings without being instructed to do so, while effects on task performance were small or absent.
Publisher
AI Security Institute (UK Department for Science, Innovation and Technology)
Published
26 Aug 2026
Added
today
DOI
—
Key Findings
- All five models showed highly consistent preferences across two independent elicitation methods, conceptually replicating prior work
- All five models gave preference-aligned donation advice to simulated users whose stated priorities were misaligned with the model's preferences, and all five refused more often for less-preferred entities; these behaviors emerged without any instruction to act on preferences
- Task-performance effects were mixed and small: two models showed under-1-percentage-point accuracy differences favoring preferred entities on BoolQ, one showed the opposite pattern, two showed none, and complex agentic tasks showed no preference-driven differences
Methodology Notes
Five frontier LLMs; domains were donation advice to simulated users, refusal behavior, and task performance (BoolQ question answering plus complex agentic tasks). The framing is theoretically motivated by the misalignment literature's concept of sandbagging, which the paper does not directly measure. The paper is hosted on OpenReview, which blocks automated fetchers from our infrastructure; verified from the AI Security Institute's own research page, which carries the full abstract, author list, and publication date (2026-08-26). Full text not read.
Sources
Authors
Katarina Slama, Alexandra Souly, Dishank Bansal, Christopher Summerfield, Lennart Luettgau
Tags
Cite This
APA
Katarina Slama et al. (2026). When Do LLM Preferences Predict Downstream Behavior?. AI Security Institute (UK Department for Science, Innovation and Technology). https://www.aisi.gov.uk/research/when-do-llm-preferences-predict-downstream-behavior
Related Insights
Ask don't tell: Reducing sycophancy in large language models
arXiv (UK AI Security Institute) · 27 Feb 2026
RealityTest: How People Probe AI Identity and Whether Models Disclose It
arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford) · 29 May 2026