26 artifacts matching
Lab publication
System Card: Claude Haiku 5.5
System card for Claude Haiku 5.5, Anthropic's small-model release of October 2026, condensed relative to its frontier cards. It reports Responsible Scaling Policy and cyber evaluations, harmlessness…
Lab publication
System Card: Claude Sonnet 5.5
System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Lab publication
Detecting and countering misuse of AI: September 2026
Anthropic's periodic threat-intelligence report covering December 2025 to August 2026 across seven harm areas. One case study, GTG-15001, documents a China-based app studio that used Claude to build…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Benchmark / dataset
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…
Lab publication
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…
Lab publication
Enabling independent research on how people use Claude
A report on a pilot in which three external research groups designed their own studies and ran them through Anthropic Insights, the company's privacy-preserving aggregate analysis tool, on roughly 25…
Lab publication
Funding better evaluations of AI's impact on wellbeing
Announcement of a $5 million grant programme funding independent research into how AI affects user wellbeing, offering money, model access and technical support to teams building open-source evaluati…
Lab publication
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Lab publication
Claude's values across models and languages
An observational study of 309,815 anonymised production conversations characterising the values an assistant expresses and how that expression varies by model version and by the user's language. Expr…
Lab publication
System Card: Claude Sonnet 5
Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…
Lab publication
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
Lab publication
System Card: Claude Opus 4.8
Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…
Preprint
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
Lab publication
How people ask Claude for personal guidance
An Anthropic research analysis of roughly 38,000 personal-guidance conversations (sampled from about 1M) covering significant life decisions across health/wellness, career, relationships, and persona…
Lab publication
Emotion Concepts and their Function in a Large Language Model
Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…
NGO report
Claude: AI Risk Assessment
A structured product risk assessment of Claude's default chat mode and Health mode (beta) on web and mobile, free and Pro, using test accounts representing ages 13 to 17 and scored against the Instit…
Lab publication
What 81,000 People Want from AI
Anthropic report on 80,508 open-ended interviews that an AI interviewer (Anthropic Interviewer, a prompted version of Claude) conducted with Claude.ai users in 159 countries and 70 languages during o…
Peer-reviewed
The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedi…