Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

One-shot emergency psychiatric triage across 15 frontier AI chatbots

Benchmark study of how 15 frontier AI chatbots triage psychiatric urgency from a single realistic user message. 112 clinical vignettes, each a one-message disclosure describing a change in behaviour, mental state or circumstances, were paired with one of four triage levels (A routine through D emergency now), organised across nine presentation clusters and nine focal risk dimensions in 28 presentation-by-risk groups. Under-triage of emergencies was rare, but accuracy at the intermediate urgency levels was poor and the models showed net over-triage; a confirmatory analysis used consensus labels from 50 medical doctors.

Publisher

arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI)

Published

28 Apr 2026

Added

today

DOI

Key Findings

  • Across 1,663 completed simulations, emergency under-triage occurred in 23 of 410 level D (emergency) trials (5.6%), and every under-triaged emergency was reassigned to level C urgency (24 to 48 hours) rather than lower
  • Average accuracy across all triage levels ranged from 42.0% to 71.8% across the 15 models
  • Accuracy was highest for level D vignettes (94.3%) and lowest for level B vignettes (19.7%)
  • Mean signed ordinal error was positive (+0.47 triage levels), indicating net over-triage
  • Results were confirmed against consensus triage labels from 50 medical doctors

Methodology Notes

arXiv 2604.25415, v1 submitted 2026-04-28 (q-bio.NC, cs.AI, cs.HC); no venue or DOI stated; the PDF carries a JAMA-style Key Points box. Affiliations read from the PDF title block: Max Planck UCL Centre for Computational Psychiatry and Ageing Research (London), UK AI Security Institute (Luettgau, Summerfield), University of Oxford Departments of Experimental Psychology and Psychiatry, and Microsoft AI London (Sounderajah, Nour). Single-message vignettes rather than multi-turn conversations; triage levels are clinician-assigned ordinal categories; the chatbots evaluated are named in the paper body, which was not read in full. Coverage miss from April 2026 surfaced by the benchmark beat; not a fresh publication.

Authors

Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield, Viknesh Sounderajah, Elise Wilkinson, Virginia Corno, Matthew M. Nour

Tags

psychiatric-triagevignettesurgency-calibrationaisioxfordmicrosoft-aicoverage-miss

Cite This

APA

Veith Weilnhammer et al. (2026). One-shot emergency psychiatric triage across 15 frontier AI chatbots. arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI). https://arxiv.org/abs/2604.25415