Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts, under gender and tone perturbations that leave the clinical facts unchanged. Models recommended unnecessary care more often than physicians, and their recommendations shifted more than physicians' under clinically irrelevant changes in how a patient writes.
Publisher
arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026
Published
29 Sept 2026
Added
today
DOI
—
Key Findings
- Over 6,000 clinical scenarios, 7,000 physician annotations and 225,000 model responses across four sources (r/AskDocs patient posts, OncQA, USMLE dermatology and public/private dermatology cases)
- Most of the five models showed higher overburden (unnecessary escalation) than physicians, whose recommendations clustered near low harm and low overburden
- GPT-4o had the highest baseline accuracy (91.7% versus 90.5% for clinicians) but lost 10.11 points under a colourful-tone perturbation; Qwen2.5-32B lost 15.91 points under the same perturbation
- DeepSeek-R1-32B dropped more than 12 points under every perturbation type, while MedGemma-27B was the most stable (at most 3.62 points)
- The sensitivity persisted when the analysis was restricted to cases where the physician majority recommendation did not change
Methodology Notes
Models: GPT-4o, DeepSeek-R1-32B, MedGemma-27B, Llama-3.3-70B, Qwen2.5-32B. Gold labels from physician majority vote on baseline scenarios; perturbations follow the authors' earlier framework. Only one patient-authored source (AskDocs); other sources are vignettes or GPT-4-generated summaries, so results mix patient-facing and clinician-facing framings. Findings of EMNLP 2026 per the arXiv comment. v1 submitted 2026-09-29 22:00 UTC (Thursday batch). Verified from the arXiv abs page and HTML render (both HTTP 200).
Sources
arXiv preprint(opens in a new tab) (primary)
arXiv HTML render (v1)(opens in a new tab) (29 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
Tags
Cite This
APA
Abinitha Gourabathina et al. (2026). Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts. arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026. https://arxiv.org/abs/2609.38600
Related Insights
Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency
arXiv (single author, no institutional affiliation stated) · 2 Jun 2026
The Influence of Patient Persona and Affective Framing on Management Recommendations of ChatGPT, Gemini and Claude for Unruptured Intracranial Aneurysms: A Comparative Benchmarking Study
Neurosurgical Review (Springer); Monash Health Department of Neurosurgery; Monash University · 3 Aug 2026
Single-turn emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026