Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

13 artifacts matching

24 Jul 2026 Anthropic Lab publication

System Card: Claude Opus 5

System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…

30 Jun 2026 Anthropic Lab publication

System Card: Claude Sonnet 5

Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…

9 Jun 2026 Anthropic Lab publication

System Card: Claude Fable 5 & Claude Mythos 5

Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…

28 May 2026 Anthropic Lab publication

System Card: Claude Opus 4.8

Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…

1 May 2026 Harvard Business School Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…

30 Apr 2026 Anthropic Lab publication

How people ask Claude for personal guidance

An Anthropic research analysis of roughly 38,000 personal-guidance conversations (sampled from about 1M) covering significant life decisions across health/wellness, career, relationships, and persona…

9 Apr 2026 Anthropic (Transformer Circuits) Lab publication

Emotion Concepts and their Function in a Large Language Model

Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…

1 Feb 2026 Anthropic Lab publication

System Card: Claude Opus 4.6

Anthropic's 213-page system card for Claude Opus 4.6, notable for an expanded 'user wellbeing evaluations' section covering child safety, suicide and self-harm, and eating disorders, alongside sycoph…

27 Jan 2026 Anthropic Preprint

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

Empirical study of how assistant interactions affect human autonomy, based on analysis of 1.5 million real consumer conversations with Claude. Severe disempowerment-risk patterns appear in fewer than…

18 Dec 2025 Anthropic Lab publication

Protecting the wellbeing of our users

Anthropic describes its methodology and results for evaluating and improving Claude's handling of mental-health-crisis conversations, covering synthetic safety evaluations, 'prefill' stress-testing o…

27 Jun 2025 Anthropic Lab publication

How people use Claude for support, advice, and companionship

Anthropic's first large-scale study of 'affective use' of Claude, analyzing how people turn to the model for emotional support, advice, and companionship. Using the privacy-preserving Clio analysis t…

8 Jun 2024 Anthropic Lab publication

Claude's Character

An Anthropic research post describing 'character training', the alignment fine-tuning stage first applied to Claude 3 to cultivate broad dispositional traits such as curiosity, open-mindedness, and h…

20 Oct 2023 Anthropic Lab publication

Towards Understanding Sycophancy in Language Models

Demonstrates that five state-of-the-art AI assistants consistently exhibit sycophancy — matching a user's stated belief over the truthful answer — across varied free-form tasks. Traces the behaviour…