Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Benchmark testing whether language models identify the right urgency of care from user medical presentations. It harmonises five public datasets (user conversations, forum posts, clinical vignettes, patient portal messages) under a four-level acuity scale from home monitoring to immediate emergency care. Twelve models are scored in a multiple-choice format and in free-form conversational replies graded by a rubric judge, plus a physician-labelled ambiguous split.

Publisher

arXiv (Columbia University; Columbia University Irving Medical Center)

Published

12 May 2026

Added

today

DOI

—

Key Findings

  • 914 cases: 697 consensus cases and 217 physician-confirmed ambiguous cases; four acuity levels
  • Clear-case exact-match accuracy in QA format ranged from 53.3% (Llama 3.3 70B) to 85.3% (Claude Opus 4.7), mean 70.6% across 12 models
  • Conversational replies reduced over-triage but increased under-triage relative to QA, especially for higher-acuity cases
  • Lowest-acuity (A-level) error rates ranged from 13% to 84.4% across models
  • On ambiguous cases no model matched the spread of physician judgments; model predictions were more concentrated than expert uncertainty

Methodology Notes

Five public source datasets relabelled to a shared four-level framework by emergency physicians; five samples per case, modal label with ties toward higher acuity; conversational judge validated against two blinded human annotators (judge-human agreement 74.2% vs human-human 64.4% over 132 responses). Models: GPT-5.4, GPT-5-mini, GPT-4.1, Claude Opus 4.7, Sonnet 4.6, Haiku 4.5, Gemini 2.5 Pro and Flash, plus open-weight Llama 3.3 70B, Qwen 2.5 72B and others. v1 posted 2026-05-12; v2 2026-09-25 (41 pages) states it is under review for the NeurIPS 2026 Evaluations and Datasets track. Verified from arXiv abs (200) and HTML render (200).

Authors

Robin Linzmayer, Georgianna Lin, Di Coneybeare, Jason Chu, Trudi Cloyd, Manish Garg, Miles Gordon, Elizabeth Hartofilis, Benjamin Hong, Ashraf Hussain, Eugene Y. Kim, Oluchi Iheagwara King, Ross McCormack, Erica Olsen, John K. Riggins Jr., Mustafa N. Rasheed, Dana L. Sacco, Vinay Saggar, Osman R. Sayan, Amit Shembekar, Janice Shin-Kim, Wendy W. Sun, Bernard P. Chang, David Kessler, Noémie Elhadad

Tags

triageacuityunder-triageconsumer-healthneurips-2026-submission

Cite This

APA

Robin Linzmayer et al. (2026). AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment. arXiv (Columbia University; Columbia University Irving Medical Center). https://arxiv.org/abs/2605.11398