Skip to main content

Browse the library

The complete record — 658 artifacts, last updated 8 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

26 artifacts matching

7 Oct 2026 Anthropic Lab publication

Lab publication

System Card: Claude Haiku 5.5

System card for Claude Haiku 5.5, Anthropic's small-model release of October 2026, condensed relative to its frontier cards. It reports Responsible Scaling Policy and cyber evaluations, harmlessness…

28 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Sonnet 5.5

System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…

22 Sept 2026 Anthropic Lab publication

Lab publication

Claude Opus 5.5 System Card

230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…

10 Sept 2026 Anthropic (Threat Intelligence) Lab publication

Lab publication

Detecting and countering misuse of AI: September 2026

Anthropic's periodic threat-intelligence report covering December 2025 to August 2026 across seven harm areas. One case study, GTG-15001, documents a China-based app studio that used Claude to build…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

31 Aug 2026 Transluce Benchmark / dataset

Benchmark / dataset

Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)

Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…

28 Aug 2026 Anthropic Lab publication

Lab publication

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…

26 Aug 2026 Anthropic (Societal Impacts), with Stanford SALT Lab, Oxford Human Information Processing Lab and METR Lab publication

Lab publication

Enabling independent research on how people use Claude

A report on a pilot in which three external research groups designed their own studies and ran them through Anthropic Insights, the company's privacy-preserving aggregate analysis tool, on roughly 25…

25 Aug 2026 Anthropic Lab publication

Lab publication

Funding better evaluations of AI's impact on wellbeing

Announcement of a $5 million grant programme funding independent research into how AI affects user wellbeing, offering money, model access and technical support to teams building open-source evaluati…

24 Jul 2026 Anthropic Lab publication

Lab publication

System Card: Claude Opus 5

System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…

13 Jul 2026 Anthropic Lab publication

Lab publication

Claude's values across models and languages

An observational study of 309,815 anonymised production conversations characterising the values an assistant expresses and how that expression varies by model version and by the user's language. Expr…

30 Jun 2026 Anthropic Lab publication

Lab publication

System Card: Claude Sonnet 5

Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…

9 Jun 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5 & Claude Mythos 5

Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…

28 May 2026 Anthropic Lab publication

Lab publication

System Card: Claude Opus 4.8

Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…

1 May 2026 Harvard Business School Preprint

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…

30 Apr 2026 Anthropic Lab publication

Lab publication

How people ask Claude for personal guidance

An Anthropic research analysis of roughly 38,000 personal-guidance conversations (sampled from about 1M) covering significant life decisions across health/wellness, career, relationships, and persona…

9 Apr 2026 Anthropic (Transformer Circuits) Lab publication

Lab publication

Emotion Concepts and their Function in a Large Language Model

Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…

25 Mar 2026 Youth AI Safety Institute, Common Sense Media NGO report

NGO report

Claude: AI Risk Assessment

A structured product risk assessment of Claude's default chat mode and Health mode (beta) on web and mobile, free and Pro, using test accounts representing ages 13 to 17 and scored against the Instit…

18 Mar 2026 Anthropic Lab publication

Lab publication

What 81,000 People Want from AI

Anthropic report on 80,508 open-ended interviews that an AI interviewer (Anthropic Interviewer, a prompted version of Claude) conducted with Claude.ai users in 159 countries and 70 languages during o…

1 Mar 2026 Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026) Peer-reviewed

Peer-reviewed

The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations

Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedi…