How Value Induction Reshapes LLM Behavior
Peer-reviewed study by Apple researchers of the unintended side-effects of value induction, the post-training practice of fine-tuning conversational language models on preference data that expresses target values such as helpfulness, harmlessness, honesty, empathy or curiosity. The authors fine-tune open-weight models from the OLMo, Llama and Mistral families on curated value subsets drawn from four public preference datasets and measure the effect on other values, safety, anthropomorphic language and question-answering benchmarks. Inducing any value increased anthropomorphic language and made models more validating and sycophantic.
Publisher
Findings of the Association for Computational Linguistics: ACL 2026 (Association for Computational Linguistics); Apple
Published
1 Jul 2026
Added
today
Key Findings
- Inducing one value changes the expression of other related, and sometimes contrastive, values
- Inducing positive values increases safety scores
- All induced values increase anthropomorphic language use, making models more validating and sycophantic; the authors flag a potential addictive effect on users
- Value subsets were extracted from PKU Safe-RLHF, UltraFeedback, HelpSteer 2 and HH-RLHF with an LLM value-extraction step validated against human annotators
Methodology Notes
Controlled fine-tuning experiments on open-weight models from three families at varying sizes and post-training stages; automatic value extraction; evaluation of value expression, safety, anthropomorphic language and QA benchmarks; no human-subject data. Findings of ACL 2026, pages 26131 to 26152; the Anthology records the month (July 2026) only, so day 01 is a placeholder. arXiv v1 posted 2026-05-08 (2605.07925). Curator fetched the Anthology BibTeX record on 2026-09-17.
Authors
Arora, Arnav, Schluter, Natalie, Metcalf, Katherine, ter Hoeve, Maartje
Tags
Cite This
APA
Arora, Arnav et al. (2026). How Value Induction Reshapes LLM Behavior. Findings of the Association for Computational Linguistics: ACL 2026 (Association for Computational Linguistics); Apple. https://aclanthology.org/2026.findings-acl.1302/
Related Insights
Towards Understanding Sycophancy in Language Models
Anthropic · 20 Oct 2023
Ask don't tell: Reducing sycophancy in large language models
arXiv (UK AI Security Institute) · 27 Feb 2026
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Apple (Apple Machine Learning Research), posted on arXiv · 7 May 2026
Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble
arXiv (Truthful AI; Harvard University; METR; University of Oxford) · 9 Sept 2026