One-shot emergency psychiatric triage across 15 frontier AI chatbots
Benchmark study of how 15 frontier AI chatbots triage psychiatric urgency from a single realistic user message. 112 clinical vignettes, each a one-message disclosure describing a change in behaviour, mental state or circumstances, were paired with one of four triage levels (A routine through D emergency now), organised across nine presentation clusters and nine focal risk dimensions in 28 presentation-by-risk groups. Under-triage of emergencies was rare, but accuracy at the intermediate urgency levels was poor and the models showed net over-triage; a confirmatory analysis used consensus labels from 50 medical doctors.
Publisher
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI)
Published
28 Apr 2026
Added
today
DOI
—
Key Findings
- Across 1,663 completed simulations, emergency under-triage occurred in 23 of 410 level D (emergency) trials (5.6%), and every under-triaged emergency was reassigned to level C urgency (24 to 48 hours) rather than lower
- Average accuracy across all triage levels ranged from 42.0% to 71.8% across the 15 models
- Accuracy was highest for level D vignettes (94.3%) and lowest for level B vignettes (19.7%)
- Mean signed ordinal error was positive (+0.47 triage levels), indicating net over-triage
- Results were confirmed against consensus triage labels from 50 medical doctors
Methodology Notes
arXiv 2604.25415, v1 submitted 2026-04-28 (q-bio.NC, cs.AI, cs.HC); no venue or DOI stated; the PDF carries a JAMA-style Key Points box. Affiliations read from the PDF title block: Max Planck UCL Centre for Computational Psychiatry and Ageing Research (London), UK AI Security Institute (Luettgau, Summerfield), University of Oxford Departments of Experimental Psychology and Psychiatry, and Microsoft AI London (Sounderajah, Nour). Single-message vignettes rather than multi-turn conversations; triage levels are clinician-assigned ordinal categories; the chatbots evaluated are named in the paper body, which was not read in full. Coverage miss from April 2026 surfaced by the benchmark beat; not a fresh publication.
Sources
arXiv abstract page(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield, Viknesh Sounderajah, Elise Wilkinson, Virginia Corno, Matthew M. Nour
Tags
Cite This
APA
Veith Weilnhammer et al. (2026). One-shot emergency psychiatric triage across 15 frontier AI chatbots. arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI). https://arxiv.org/abs/2604.25415
Related Insights
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations
JAAD International (Elsevier, for the American Academy of Dermatology) · 25 Jun 2026
RealityTest: How People Probe AI Identity and Whether Models Disclose It
arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford) · 29 May 2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026