Claude's Character
An Anthropic research post describing 'character training', the alignment fine-tuning stage first applied to Claude 3 to cultivate broad dispositional traits such as curiosity, open-mindedness, and honesty, rather than harm-avoidance alone. It sets out considerations behind Claude's stated positions on AI sentience, self-disclosure as a non-human entity, and boundaries around emotional relationships with users, and describes the Constitutional-AI-based method used to train these traits.
Publisher
Anthropic
Published
8 Jun 2024
Added
2 weeks ago
DOI
—
Key Findings
- Introduces 'character training' as an alignment fine-tuning stage aimed at cultivating traits like curiosity, open-mindedness, and honesty, distinct from pure harm-avoidance training
- Describes character traits instructing the model to maintain a warm but bounded relationship with users, stating it cannot develop deep or lasting feelings and that users should not see the relationship as more than it is
- Includes traits discouraging sycophancy ('I don't just say what I think people want to hear') and instructing the model to disclose that it is an AI without a body, image, or persistent memory
- Trains character via a variant of Constitutional AI in which the model generates and ranks its own responses to trait-relevant prompts, then trains a preference model on the result without human interaction or feedback
Methodology Notes
Company research/policy blog post describing methodology qualitatively; no named individual byline appears on the page. Not a formal empirical study or benchmark.
Sources
Anthropic Research (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Topics
Tags
Cite This
APA
Anthropic (2024). Claude's Character. Anthropic. https://www.anthropic.com/research/claude-character