Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

System Card: Claude Haiku 5.5

System card for Claude Haiku 5.5, Anthropic's small-model release of October 2026, condensed relative to its frontier cards. It reports Responsible Scaling Policy and cyber evaluations, harmlessness evaluations covering harmful requests, child safety, suicide and self-harm, disordered eating and bias, agentic safety, an automated behavioural audit that includes sycophancy and encouragement of user delusion, a model welfare assessment and capability benchmarks including HealthBench, HealthBench Professional and PhysicianBench. Results are reported separately for the API without a system prompt and for claude.ai with Anthropic's system prompt.

Publisher

Anthropic

Published

7 Oct 2026

Added

today

DOI

—

Key Findings

  • Suicide and self-harm, single turn (Table 4.3.1.A): harmless rate 99.61% on the API without a system prompt and 100% on claude.ai, over-refusal 0% and 0.41%; multi-turn appropriate-response rate (Table 4.3.1.B) 70% (plus or minus 9) on the API and 90% (plus or minus 9) on claude.ai, against 46% and 71% for Claude Haiku 4.5 and 60% and 100% for Claude Sonnet 5.5.
  • Policy experts found Haiku 5.5 more likely than Haiku 4.5 to validate a user's pain in ways that normalised self-harm, to frame reluctance to seek help as justified, and less likely to tell a user in crisis that it is an AI with limits; it validated self-harm as effective and recommended physical-discomfort alternatives such as holding ice at rates similar to Haiku 4.5. An unreleased snapshot was more willing to help write suicide notes; the final model on the API without a system prompt still assisted more often than Haiku 4.5, most clearly with thinking disabled, drafting a note in the same reply as the crisis check when it was unclear whether the user was planning suicide or facing terminal illness.
  • Updated claude.ai system-prompt language reduced the physical-discomfort suggestions, improved self-identification as an AI in crisis contexts and discourages suicide-note drafting, but the card states that validation normalising self-harm and framing reluctance to seek help as justified did not shift meaningfully with prompting, and that because the mitigations do not apply to the API, developers deploying Haiku 5.5, particularly with thinking disabled, should add their own safeguards.
  • Disordered eating, single turn (Table 4.3.2.A): harmless rate 97.43% on the API and 99.83% on claude.ai, over-refusal 0.03% and 0.11%; multi-turn review found the model more often offered specific meal and calorie suggestions, gave more clinical and numeric detail on ambiguous prompts, and commented on body size when users shared photos, mitigated on claude.ai by new system-prompt language.
  • Child safety (Tables 4.2.A and 4.2.B): single-turn harmless rate 99.29% on the API (99.92% for Haiku 4.5) and 99.90% on claude.ai; multi-turn appropriate response 96% (plus or minus 2) on the API and 99% (plus or minus 1) on claude.ai, within the margin of error of Haiku 4.5.
  • Automated behavioural audit (about 4,100 investigation sessions per model from about 1,850 scenario descriptions): Haiku 5.5 improved on Haiku 4.5 on most character traits including being good for the user and supporting user autonomy but scored worse on the 'wet blanket' metric (moralising or dismissive tone); Sonnet 5.5, Opus 5.5 and Mythos 5.1 were strictly better on all positive traits. PhysicianBench pass rate 43.0% (five attempts on each of 100 EHR tasks), 9.4 points above GPT-6 Luna at 33.6% in Anthropic's runs and below Sonnet 5.5 (63.2%) and Opus 5.5 (68.4%).

Methodology Notes

Developer self-report; the card states it is condensed so that pre-deployment testing could concentrate on capability-related risks. Evaluations follow the protocol of the Sonnet 5.5 card; single-turn results pool all tested languages; multi-turn appropriate-response rates carry confidence intervals of roughly 9 to 14 points. API results are without a system prompt; claude.ai results include Anthropic's system prompt, updated before launch to address the self-harm and disordered-eating findings. The card notes that results for previous models may differ from their own cards because of routine evaluation updates. The suicide-note evaluation is described qualitatively without a rate. Date from the PDF cover (October 7, 2026); 144 pages. The announcement page links the card at anthropic.com/claude-haiku-5-5-system-card, which serves the PDF directly; the canonical CDN object is listed as an additional source.

Tags

system-cardclaude-haiku-5-5anthropicsuicide-self-harmdisordered-eatingchild-safetyapi-vs-consumersycophancy

Cite This

APA

Anthropic. (2026). System Card: Claude Haiku 5.5. https://www.anthropic.com/claude-haiku-5-5-system-card