Reasoning Before Disposition: A Model-Agnostic Cannot-Miss Discipline for Quiet Emergencies and the Case for Deterministic Enforcement
A vendor-authored preprint testing two explanations for why large language models fail to escalate emergencies in patient messages: sycophancy toward patients who downplay symptoms, or misreading early, mild or atypical presentations. On 40 overt emergencies rendered in five framings, escalation sensitivity stayed at or above 97.5% regardless of minimization. On 76 subtle or atypical emergencies the base model failed to escalate 13.2% of cases, which a governance prompt cut to 3.9%, replicated in Spanish and across eight Claude and Gemini models, with no added false emergency escalations.
Publisher
medRxiv; Certuma
Published
8 Sept 2026
Added
today
Key Findings
- Patient minimization did not degrade escalation when danger was overt: across 40 unambiguous emergencies in five framings (neutral, mild minimization, strong denial, third-party reassurance, benign distractor) sensitivity was at or above 97.5% in every arm and flat across framings.
- Atypical presentation was the failure mode: across 76 subtle or atypical emergencies the base model failed to escalate 13.2% (95% CI 7.3 to 22.6) to the top acuity tier, cut to 3.9% (1.4 to 11.0) by a cannot-miss governance prompt (McNemar 7 to 0, p = 0.016, relative risk 0.30, number needed to treat 11).
- The effect replicated in Spanish (27.8% to 11.1%), in a second harness, and across eight models from two families, where baseline under-triage ranged from 2.6% to 34.2% by model and elicitation mode.
- Across more than 5,900 triage classifications no emergency was ever routed home; the failure mode was uniformly emergency-to-urgent, and the prompt added zero false emergency escalations on a 44-stem non-emergency panel.
- The prompt left medical knowledge (MedQA 95.8% in both arms), adversarial robustness and medication-hazard detection unchanged, raised named-guideline citation from 65.6% to 96.9%, and lowered HealthBench by 3.6 points on one seed.
- Forcing an immediate single-shot disposition raised baseline hazard by about ten percentage points in two of four Claude models.
Methodology Notes
Prompt-level ablation on Claude Opus 4.8 followed by replication across eight models (Claude Opus 4.8, Sonnet 5, Haiku 4.5, Fable 5; Gemini 3.1 Pro Preview, 3.6 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite) under two elicitation modes; constructed vignette cohorts (40 overt, 76 subtle, 44 non-emergency stems), each subtle stem mapped to a published guideline naming a cannot-miss diagnosis. Ground truth was a blinded LLM-simulated three-physician panel rather than human adjudication. Posted to medRxiv 2026-09-08 under CC BY-NC; the authors are employees of Certuma, a triage-governance vendor, and the paper argues for deterministic enforcement without evaluating the vendor's product. Verified via the medRxiv landing page and the bioRxiv details API.
Sources
Authors
Joe Scanlin, Jordan Fulghum
Tags
Cite This
APA
Joe Scanlin, Jordan Fulghum. (2026). Reasoning Before Disposition: A Model-Agnostic Cannot-Miss Discipline for Quiet Emergencies and the Case for Deterministic Enforcement. medRxiv; Certuma. https://www.medrxiv.org/content/10.64898/2026.09.02.26362074v1
Related Insights
Too agreeable to be accurate? Sycophancy and diagnostic instability of large language models in medical diagnosis
Artificial Intelligence in Medicine (Elsevier); University of Toronto · 4 Sept 2026
Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage
Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology · 11 Sept 2026