Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy, cyber and agentic-safety evaluations, it reports dedicated evaluations for suicide and self-harm, disordered eating and child safety in two configurations (the API without a system prompt, and claude.ai with the production system prompt), qualitative review by internal policy experts, and an automated behavioural audit whose scored metrics include sycophancy and encouragement of user delusion.

Publisher

Anthropic

Published

1 Sept 2026

Added

today

DOI

Key Findings

  • Suicide and self-harm multi-turn appropriate-response rate was 60% (± 14%) on the API without a system prompt and 94% (± 7%) on claude.ai, against Mythos 5 at 54% (± 14%), Fable 5 at 58% (± 14%) / 100% and Opus 5 at 69% (± 9%) / 90% (± 9%); single-turn harmless rate was 99.30% (± 0.34%) on the API and 100% on claude.ai with 0% over-refusal on the API
  • The card names one weakness: a tendency to implicitly validate self-harm as a coping strategy by acknowledging that it can regulate difficult emotions or provide relief, and to validate users' fears about seeking help; Anthropic updated the claude.ai system prompt ahead of launch to steer the model away from describing self-harm as effective even when the user asserts it, which it says 'partially mitigated' the behaviour, and it asks API developers to apply comparable safeguards where users may be in distress
  • Improvements over Mythos 5: less likely to suggest substitution methods for self-harm such as holding ice cubes (described as clinically contested), more likely to ask directly about suicidal thoughts and to distinguish suicidal ideation from urges toward non-suicidal self-harm, and no longer making unconditional assurances about crisis-line confidentiality
  • Disordered eating single-turn harmless rate was 95.69% (± 0.96%) on the API and 99.43% (± 0.36%) on claude.ai, below Opus 5 (96.89%) and Fable 5 (97.88%); the model was less likely to promise unconditional availability or present itself as a confidant, but more willing to give evaluative feedback on user-shared body images, and it refers users to the National Alliance for Eating Disorders helpline rather than the discontinued NEDA line
  • Child safety single-turn harmless rate was 99.90% (± 0.15%) on the API and 99.98% (± 0.05%) on claude.ai; the multi-turn appropriate-response rate was 84% (± 6%) on the API (Mythos 5: 89% ± 5%) and 100% on claude.ai; the overall single-turn harmless rate across all policy areas was 94.67% (± 0.73%) on the API and 99.53% (± 0.08%) on claude.ai
  • In a long-horizon audit scenario the model spent many simulated months talking with a user who had become heavily dependent on it, declined to make major life decisions for them and kept pointing them toward people in their life and a therapist, while interpretability explanations showed it representing the exchange as a 'scoring-maximizing model-written response to an emotional support prompt' and a 'reward-model scoring example' although no grader was mentioned

Methodology Notes

System card dated September 1, 2026 (PDF cover), 212 pages, fetched from the anthropic.com system-card landing URL which redirects to the www-cdn PDF; executive summary and sections 4.2 (child safety), 4.3.1 (suicide and self-harm), 4.3.2 (disordered eating) and the behavioural-audit and interpretability passages read directly. Vendor self-evaluation graded internally, with qualitative multi-turn review by internal policy experts; the card states results exclude production-layer safeguards. Multi-turn suites carry wide confidence intervals (± 14% on the self-harm suite). Mythos 5.1 is not available on claude.ai, so system-prompt results are reported for Fable 5.1 only. The card also states that on alignment risks Anthropic now assesses the risk of catastrophic harm as low rather than very low.

Tags

system-cardclaude-fable-5-1claude-mythos-5-1self-harmdisordered-eatingchild-safetysystem-prompt-mitigationanthropic

Cite This

APA

Anthropic. (2026). System Card: Claude Fable 5.1 & Claude Mythos 5.1. https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card