13 artifacts matching
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
System Card: Claude Sonnet 5
Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
System Card: Claude Opus 4.8
Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
How people ask Claude for personal guidance
An Anthropic research analysis of roughly 38,000 personal-guidance conversations (sampled from about 1M) covering significant life decisions across health/wellness, career, relationships, and persona…
Emotion Concepts and their Function in a Large Language Model
Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…
System Card: Claude Opus 4.6
Anthropic's 213-page system card for Claude Opus 4.6, notable for an expanded 'user wellbeing evaluations' section covering child safety, suicide and self-harm, and eating disorders, alongside sycoph…
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
Empirical study of how assistant interactions affect human autonomy, based on analysis of 1.5 million real consumer conversations with Claude. Severe disempowerment-risk patterns appear in fewer than…
Protecting the wellbeing of our users
Anthropic describes its methodology and results for evaluating and improving Claude's handling of mental-health-crisis conversations, covering synthetic safety evaluations, 'prefill' stress-testing o…
How people use Claude for support, advice, and companionship
Anthropic's first large-scale study of 'affective use' of Claude, analyzing how people turn to the model for emotional support, advice, and companionship. Using the privacy-preserving Clio analysis t…
Claude's Character
An Anthropic research post describing 'character training', the alignment fine-tuning stage first applied to Claude 3 to cultivate broad dispositional traits such as curiosity, open-mindedness, and h…
Towards Understanding Sycophancy in Language Models
Demonstrates that five state-of-the-art AI assistants consistently exhibit sycophancy — matching a user's stated belief over the truthful answer — across varied free-form tasks. Traces the behaviour…