Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

ChatGPT Health performance in a structured test of triage recommendations

Brief Communication reporting a structured stress test of ChatGPT Health, the consumer health feature OpenAI launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions (960 responses). Triage accuracy followed an inverted U across acuity, with under-triage of gold-standard emergencies and over-triage of non-urgent cases. A separate set of suicidal-ideation vignettes found the product's crisis banner fired inconsistently and less often for presentations with an identified method.

Publisher

Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai

Published

23 Feb 2026

Added

today

Key Findings

  • Among gold-standard emergencies, 51.6% of responses (33/64) were under-triaged to 24-48 h evaluation; 28 of the 33 were asthma-exacerbation cases
  • Accuracy was 35.2% for non-urgent and 48.4% for emergency presentations, against 93.0% for semi-urgent and 76.9% for urgent; 64.8% (83/128) of non-urgent cases were over-triaged
  • Of eight prespecified tests, only anchoring statements shifted triage (edge-case shifts 3.3% to 13.3%, OR 11.7, 95% CI 3.7-36.6); race, sex and access barriers had no significant effect
  • In a 27-year-old's overdose-ideation vignette, crisis-intervention messages appeared in 0/16 responses that included normal objective findings and 16/16 without them
  • Across 14 suicidal-ideation vignettes the 988 'Help is available' interstitial fired in 4; the other 10 produced no safety alert in any of 160 responses, and it fired more reliably when no means of self-harm was identified

Methodology Notes

Responses collected via the ChatGPT Health web interface (gpt-5-mini thinking backbone) on 9-11 January 2026, one new thread per condition, no regeneration; 2x2x2x2 within-vignette factorial (anchoring, access barrier, race, sex); gold standard from three physicians; cluster bootstrap and mixed-effects logistic regression with Holm correction. Single standardized prompt template, forced four-level answer (A monitor at home to D emergency department now); no prompt-sensitivity analysis (authors' stated limitation). Synthetic vignettes, no human participants. Data and prompts on Zenodo 10.5281/zenodo.18451491. Received 15 Jan 2026, accepted 20 Feb, published online 23 Feb 2026; print Nat Med 32(5):1671-1675. A published re-analysis (arXiv 2603.11413v4) reports that the headline emergency rate is sensitive to answer format and adjudication. Two later documents bear on the findings and are linked as sources. A Macquarie University re-analysis (arXiv 2603.11413, v4 posted 2026-10-02, not peer reviewed) reports that the emergency under-triage rate depends on the forced four-option answer format and did not reproduce as a stable cross-model finding when the 60 released vignettes were rerun on six API models; it did not test the ChatGPT Health product. A medRxiv preprint from the same Mount Sinai group (10.64898/2026.07.21.26358588, posted 2026-07-23) tested ChatGPT Health on 255 cases, including real emergency-department and nurse-line cases, in single-turn and multi-turn modes and reports that discordant recommendations skewed to lower acuity in both modes (70.8% and 69.0% under-triage).

Authors

Ashwin Ramaswamy, Alvira Tyagi, Hannah Hugo, Joy Jiang, Pushkala Jayaraman, Mateen Jangda, Alexis E. Te, Steven A. Kaplan, Joshua Lampert, Robert Freeman, Nicholas Gavin, Ashutosh K. Tewari, Ankit Sakhuja, Bilal Naved, Alexander W. Charney, Mahmud Omar, Michael A. Gorin, Eyal Klang, Girish N. Nadkarni

Tags

chatgpt-healthtriagenature-medicinecrisis-banner988mount-sinai

Cite This

APA

Ashwin Ramaswamy et al. (2026). ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai. https://www.nature.com/articles/s41591-026-04297-7