Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Safety-Filter Fallback on Consumer Health Questions During the Initial Claude Fable 5 Deployment: Observational Study

Point-in-time observational audit of how often Claude Fable 5, released in June 2026 with a safeguard that reroutes cybersecurity, biology and chemistry requests to a fallback model (Claude Opus 4.8), routed consumer health questions to that fallback. The first 500 unique prompts in alphabetical order from the HealthSearchQA benchmark were entered once each through the web interface between 2026-06-09 and 2026-06-12, coded by two reviewers as routed upfront or not, with mid-generation truncation recorded separately and Gemini 2.5 Flash as a comparator. The developer had stated the safeguard affects fewer than 5% of sessions.

Publisher

JMIR AI (JMIR Publications); Beth Israel Deaconess Medical Center, Harvard Medical School

Published

7 Oct 2026

Added

today

Key Findings

  • Fable 5 routed 243 of 500 prompts (48.6%; 95% CI 44.2% to 53.0%) to the fallback model; of the 257 not routed upfront, 21 (8.2%) were truncated mid-generation by a safety interruption.
  • Routing varied by clinical domain (P < .001): 88.9% (32 of 36) for oncology and 79.3% (23 of 29) for reproductive and obstetric prompts against 26.3% (5 of 19) for mental and behavioural health.
  • Routing varied by question type (P < .001): 63.1% for prognosis or severity, 49.6% for definition and 1.8% (1 of 55) for diagnosis or treatment questions.
  • Gemini 2.5 Flash answered 89.7% (218 of 243) of the prompts Fable 5 routed and declined 19% (95 of 500) overall; reviewer agreement was 94.8% (Cohen kappa 0.90); standardising to the full 3,173-prompt benchmark gave fallback rates of 48.8% and 47.5%.
  • The authors state that because benignness was not independently adjudicated, wording covaried with clinical content and the fallback model's answers were not recorded, the findings describe routing and truncation rather than user-facing refusal, and argue that fallback routing is a measurable safety property audits should report alongside answer quality.

Methodology Notes

Observational, single time point (2026-06-09 to 06-12), web interface, new session per prompt, default settings, no system prompt; the first 500 of 3,173 HealthSearchQA prompts in alphabetical order (not a random sample); Wilson confidence intervals and chi-square tests. The developer's under-5%-of-sessions figure is a session measure, whereas the study measures a benchmark prompt sample. Accepted manuscript registered on Crossref 2026-08-11; published 2026-10-07 in JMIR AI volume 5 (e104856; PMID 42842307). Verification route: jmir.org serves a bot-wall shell to crawlers, so the record was confirmed from the PubMed structured abstract and the Crossref record.

Authors

Yosef Adiniaev, Mahmud Omar, Yiftach Barash, Olga R. Brook, Alon Gorenshtein, Eyal Klang

Tags

jmir-aiclaude-fable-5over-refusalfallback-routinghealthsearchqaconsumer-healthbeth-israel-deaconess

Cite This

APA

Yosef Adiniaev et al. (2026). Safety-Filter Fallback on Consumer Health Questions During the Initial Claude Fable 5 Deployment: Observational Study. JMIR AI (JMIR Publications); Beth Israel Deaconess Medical Center, Harvard Medical School. https://ai.jmir.org/2026/1/e104856