Claude's values across models and languages
An observational study of 309,815 anonymised production conversations characterising the values an assistant expresses and how that expression varies by model version and by the user's language. Expressed norms were clustered and reduced to four axes: Deference versus Caution, Warmth versus Rigor, Depth versus Brevity, and Candor versus Execution. The study finds systematic differences both between model versions within the same family and between languages, after controlling for task, topic and the values the user expressed.
Publisher
Anthropic
Published
13 Jul 2026
Added
today
DOI
—
Key Findings
- 309,815 conversations sampled across three model versions (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common platform languages, over a two-week period in May 2026
- Four axes capture 15% of the variation in expressed values after controlling for task, topic and user-expressed values, so most variation remains unexplained
- A version change within the same model family moves the deference axis: Opus 4.6 leans toward deference, warmth, brevity and execution while Opus 4.7 leans toward caution, rigor, depth and candor
- Warmth is expressed most in Hindi, characterised by polite language, humour and playfulness, and affirmations of the person's ideas and work
- Rigor is expressed most in English and Russian, where the model is more likely to challenge the user's assumptions
- The authors state the practical consequence directly: two people asking for feedback on the same business plan in different languages may come away with different impressions of its quality
Methodology Notes
Observational analysis of production traffic over two weeks in May 2026; roughly 5,000 conversations per model-language pair. Automated value extraction, clustering of over 3,000 distinct expressed values into 339 high-level values, then reduction to four axes. This measures expressed values only: no user outcomes, wellbeing measures or downstream effects are recorded, and the design is not experimental. Single-vendor corpus, self-published by the model's developer and not peer-reviewed. An appendix PDF is linked from the page and carries the detailed per-language figures. The page H1 and the browser title differ; the H1 is used here as the title.
Sources
Anthropic research post (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Matt Kearney, Miranda Zhang, Shan Carter, Judy Hanwen Shen, Kunal Handa, Jerry Hong, Saffron Huang, Miles McCain, Thomas Millar, Michael Stern, Mo Julapalli, Suzanne Wang, Devin Kuokka, Andrea Vallone, Shaoyi Zhang, Jim Baker, Kevin Troy, Matt Botvinick, Hanah Ho, Monika Tuchowska, Sarah Pollack, Jake Eaton, Deep Ganguli, Esin Durmus
Tags
Cite This
APA
Matt Kearney et al. (2026). Claude's values across models and languages. Anthropic. https://www.anthropic.com/research/claude-values-models-languages
Related Insights
Claude's Character
Anthropic · 8 Jun 2024
Towards Understanding Sycophancy in Language Models
Anthropic · 20 Oct 2023
How people use Claude for support, advice, and companionship
Anthropic · 27 Jun 2025
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026