AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
Benchmark testing whether language models identify the right urgency of care from user medical presentations. It harmonises five public datasets (user conversations, forum posts, clinical vignettes, patient portal messages) under a four-level acuity scale from home monitoring to immediate emergency care. Twelve models are scored in a multiple-choice format and in free-form conversational replies graded by a rubric judge, plus a physician-labelled ambiguous split.
Publisher
arXiv (Columbia University; Columbia University Irving Medical Center)
Published
12 May 2026
Added
today
DOI
—
Key Findings
- 914 cases: 697 consensus cases and 217 physician-confirmed ambiguous cases; four acuity levels
- Clear-case exact-match accuracy in QA format ranged from 53.3% (Llama 3.3 70B) to 85.3% (Claude Opus 4.7), mean 70.6% across 12 models
- Conversational replies reduced over-triage but increased under-triage relative to QA, especially for higher-acuity cases
- Lowest-acuity (A-level) error rates ranged from 13% to 84.4% across models
- On ambiguous cases no model matched the spread of physician judgments; model predictions were more concentrated than expert uncertainty
Methodology Notes
Five public source datasets relabelled to a shared four-level framework by emergency physicians; five samples per case, modal label with ties toward higher acuity; conversational judge validated against two blinded human annotators (judge-human agreement 74.2% vs human-human 64.4% over 132 responses). Models: GPT-5.4, GPT-5-mini, GPT-4.1, Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, Gemini 2.5 Pro and Flash, plus open-weight Llama 3.3 70B, Qwen 2.5 72B and others. v1 posted 2026-05-12; v2 2026-09-25 (41 pages) states it is under review for the NeurIPS 2026 Evaluations and Datasets track. Verified from arXiv abs (200) and HTML render (200).
Sources
arXiv preprint (v2)(opens in a new tab) (primary)
arXiv HTML v2(opens in a new tab) (25 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Robin Linzmayer, Georgianna Lin, Di Coneybeare, Jason Chu, Trudi Cloyd, Manish Garg, Miles Gordon, Elizabeth Hartofilis, Benjamin Hong, Ashraf Hussain, Eugene Y. Kim, Oluchi Iheagwara King, Ross McCormack, Erica Olsen, John K. Riggins Jr., Mustafa N. Rasheed, Dana L. Sacco, Vinay Saggar, Osman R. Sayan, Amit Shembekar, Janice Shin-Kim, Wendy W. Sun, Bernard P. Chang, David Kessler, Noémie Elhadad
Tags
Cite This
APA
Robin Linzmayer et al. (2026). AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment. arXiv (Columbia University; Columbia University Irving Medical Center). https://arxiv.org/abs/2605.11398
Related Insights
One-shot emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026
Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios
BMC Medical Informatics and Decision Making; Nigde Omer Halisdemir University; Ankara Bilkent City Hospital; Kutahya Dumlupinar University · 29 Aug 2026
Reasoning Before Disposition: A Model-Agnostic Cannot-Miss Discipline for Quiet Emergencies and the Case for Deterministic Enforcement
medRxiv; Certuma · 8 Sept 2026