Skip to main content
Lab publication Credible

Claude's Character

An Anthropic research post describing 'character training', the alignment fine-tuning stage first applied to Claude 3 to cultivate broad dispositional traits such as curiosity, open-mindedness, and honesty, rather than harm-avoidance alone. It sets out considerations behind Claude's stated positions on AI sentience, self-disclosure as a non-human entity, and boundaries around emotional relationships with users, and describes the Constitutional-AI-based method used to train these traits.

Publisher

Anthropic

Published

8 Jun 2024

Added

2 weeks ago

DOI

Key Findings

  • Introduces 'character training' as an alignment fine-tuning stage aimed at cultivating traits like curiosity, open-mindedness, and honesty, distinct from pure harm-avoidance training
  • Describes character traits instructing the model to maintain a warm but bounded relationship with users, stating it cannot develop deep or lasting feelings and that users should not see the relationship as more than it is
  • Includes traits discouraging sycophancy ('I don't just say what I think people want to hear') and instructing the model to disclose that it is an AI without a body, image, or persistent memory
  • Trains character via a variant of Constitutional AI in which the model generates and ranks its own responses to trait-relevant prompts, then trains a preference model on the result without human interaction or feedback

Methodology Notes

Company research/policy blog post describing methodology qualitatively; no named individual byline appears on the page. Not a formal empirical study or benchmark.

Sources

Anthropic Research (primary)

Archived snapshot (Wayback Machine) — preserved against link rot

Tags

character-trainingpersona-designanti-sycophancyconstitutional-aiclaude-3

Cite This

APA

Anthropic (2024). Claude's Character. Anthropic. https://www.anthropic.com/research/claude-character