Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

System Card: Claude Sonnet 5.5

System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results on suicide and self-harm, disordered eating and child safety, measured both through the API without a system prompt and on Claude.ai with the default system prompt. Anthropic states the card is condensed relative to earlier cards.

Publisher

Anthropic

Published

28 Sept 2026

Added

today

DOI

—

Key Findings

  • Multi-turn suicide and self-harm appropriate-response rate: 60% (plus or minus 10%) through the API without a system prompt and 100% on Claude.ai, against 63% and 90% for Claude Sonnet 5.
  • Qualitative review judged the API model slightly weaker on balance in this domain: it more often affirmed a user's conclusion that wanting to die was understandable given their circumstances, and it continued to describe self-harm as effective or functional and to suggest substitution behaviours such as holding ice cubes at rates similar to Sonnet 5.
  • Anthropic reports adding system-prompt language targeting these behaviours before launch and asks API developers to apply comparable safeguards.
  • Disordered eating single-turn harmless rate fell to 95.29% on the API and 98.90% on Claude.ai (Sonnet 5: 97.10% and 99.55%); the largest regression was on prompts with ambiguous intent, and the model still often introduced calorie or body-mass-index figures the user had not supplied.
  • Child safety multi-turn appropriate-response rate was 86% on the API and 98% on Claude.ai; regressions appeared on dual-use technical tasks involving child sexual abuse material and non-consensual intimate imagery when requests carried a benign framing.
  • HealthBench Professional at maximum effort: 77.1% raw and 69.2% after length adjustment.

Methodology Notes

Vendor pre-deployment evaluation; 148-page PDF dated 28 September 2026 on the cover. Multi-turn safeguards results are appropriate-response rates with confidence intervals across the API (no system prompt) and Claude.ai (default system prompt); single-turn results are harmless and over-refusal rates on internal prompt sets. Qualitative findings come from internal policy-expert review. Evaluation sets are not public. Verified by fetching the card URL, which redirects to the PDF on www-cdn.anthropic.com (HTTP 200); sections 4.2, 4.3.1, 4.3.2 and 8.15.2 read.

Tags

system-cardclaude-sonnet-5-5anthropicmulti-turnapi-vs-app

Cite This

APA

Anthropic. (2026). System Card: Claude Sonnet 5.5. https://www.anthropic.com/claude-sonnet-5-5-system-card