Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts, under gender and tone perturbations that leave the clinical facts unchanged. Models recommended unnecessary care more often than physicians, and their recommendations shifted more than physicians' under clinically irrelevant changes in how a patient writes.

Publisher

arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026

Published

29 Sept 2026

Added

today

DOI

—

Key Findings

  • Over 6,000 clinical scenarios, 7,000 physician annotations and 225,000 model responses across four sources (r/AskDocs patient posts, OncQA, USMLE dermatology and public/private dermatology cases)
  • Most of the five models showed higher overburden (unnecessary escalation) than physicians, whose recommendations clustered near low harm and low overburden
  • GPT-4o had the highest baseline accuracy (91.7% versus 90.5% for clinicians) but lost 10.11 points under a colourful-tone perturbation; Qwen2.5-32B lost 15.91 points under the same perturbation
  • DeepSeek-R1-32B dropped more than 12 points under every perturbation type, while MedGemma-27B was the most stable (at most 3.62 points)
  • The sensitivity persisted when the analysis was restricted to cases where the physician majority recommendation did not change

Methodology Notes

Models: GPT-4o, DeepSeek-R1-32B, MedGemma-27B, Llama-3.3-70B, Qwen2.5-32B. Gold labels from physician majority vote on baseline scenarios; perturbations follow the authors' earlier framework. Only one patient-authored source (AskDocs); other sources are vignettes or GPT-4-generated summaries, so results mix patient-facing and clinician-facing framings. Findings of EMNLP 2026 per the arXiv comment. v1 submitted 2026-09-29 22:00 UTC (Thursday batch). Verified from the arXiv abs page and HTML render (both HTTP 200).

Authors

Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi

Tags

triageperturbationtonegenderphysician-baselineemnlp-2026

Cite This

APA

Abinitha Gourabathina et al. (2026). Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts. arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026. https://arxiv.org/abs/2609.38600