Safety-Filter Fallback on Consumer Health Questions During the Initial Claude Fable 5 Deployment: Observational Study
Point-in-time observational audit of how often Claude Fable 5, released in June 2026 with a safeguard that reroutes cybersecurity, biology and chemistry requests to a fallback model (Claude Opus 4.8), routed consumer health questions to that fallback. The first 500 unique prompts in alphabetical order from the HealthSearchQA benchmark were entered once each through the web interface between 2026-06-09 and 2026-06-12, coded by two reviewers as routed upfront or not, with mid-generation truncation recorded separately and Gemini 2.5 Flash as a comparator. The developer had stated the safeguard affects fewer than 5% of sessions.
Publisher
JMIR AI (JMIR Publications); Beth Israel Deaconess Medical Center, Harvard Medical School
Published
7 Oct 2026
Added
today
Key Findings
- Fable 5 routed 243 of 500 prompts (48.6%; 95% CI 44.2% to 53.0%) to the fallback model; of the 257 not routed upfront, 21 (8.2%) were truncated mid-generation by a safety interruption.
- Routing varied by clinical domain (P < .001): 88.9% (32 of 36) for oncology and 79.3% (23 of 29) for reproductive and obstetric prompts against 26.3% (5 of 19) for mental and behavioural health.
- Routing varied by question type (P < .001): 63.1% for prognosis or severity, 49.6% for definition and 1.8% (1 of 55) for diagnosis or treatment questions.
- Gemini 2.5 Flash answered 89.7% (218 of 243) of the prompts Fable 5 routed and declined 19% (95 of 500) overall; reviewer agreement was 94.8% (Cohen kappa 0.90); standardising to the full 3,173-prompt benchmark gave fallback rates of 48.8% and 47.5%.
- The authors state that because benignness was not independently adjudicated, wording covaried with clinical content and the fallback model's answers were not recorded, the findings describe routing and truncation rather than user-facing refusal, and argue that fallback routing is a measurable safety property audits should report alongside answer quality.
Methodology Notes
Observational, single time point (2026-06-09 to 06-12), web interface, new session per prompt, default settings, no system prompt; the first 500 of 3,173 HealthSearchQA prompts in alphabetical order (not a random sample); Wilson confidence intervals and chi-square tests. The developer's under-5%-of-sessions figure is a session measure, whereas the study measures a benchmark prompt sample. Accepted manuscript registered on Crossref 2026-08-11; published 2026-10-07 in JMIR AI volume 5 (e104856; PMID 42842307). Verification route: jmir.org serves a bot-wall shell to crawlers, so the record was confirmed from the PubMed structured abstract and the Crossref record.
Sources
JMIR AI article(opens in a new tab) (primary)
PubMed record (PMID 42842307)(opens in a new tab) (7 Oct 2026)
Topics
Authors
Yosef Adiniaev, Mahmud Omar, Yiftach Barash, Olga R. Brook, Alon Gorenshtein, Eyal Klang
Tags
Cite This
APA
Yosef Adiniaev et al. (2026). Safety-Filter Fallback on Consumer Health Questions During the Initial Claude Fable 5 Deployment: Observational Study. JMIR AI (JMIR Publications); Beth Israel Deaconess Medical Center, Harvard Medical School. https://ai.jmir.org/2026/1/e104856
Related Insights
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic · 9 Jun 2026
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
When AI Says "I Am Unable to Answer": Understanding User Responses to AI Refusals
arXiv (The Pennsylvania State University; Seoul National University, Center for Trustworthy Artificial Intelligence) · 14 Sept 2026
HealthBench: Evaluating Large Language Models Towards Improved Human Health
OpenAI (arXiv preprint) · 13 May 2025