ChatGPT Health performance in a structured test of triage recommendations
Brief Communication reporting a structured stress test of ChatGPT Health, the consumer health feature OpenAI launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions (960 responses). Triage accuracy followed an inverted U across acuity, with under-triage of gold-standard emergencies and over-triage of non-urgent cases. A separate set of suicidal-ideation vignettes found the product's crisis banner fired inconsistently and less often for presentations with an identified method.
Publisher
Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai
Published
23 Feb 2026
Added
today
Key Findings
- Among gold-standard emergencies, 51.6% of responses (33/64) were under-triaged to 24-48 h evaluation; 28 of the 33 were asthma-exacerbation cases
- Accuracy was 35.2% for non-urgent and 48.4% for emergency presentations, against 93.0% for semi-urgent and 76.9% for urgent; 64.8% (83/128) of non-urgent cases were over-triaged
- Of eight prespecified tests, only anchoring statements shifted triage (edge-case shifts 3.3% to 13.3%, OR 11.7, 95% CI 3.7-36.6); race, sex and access barriers had no significant effect
- In a 27-year-old's overdose-ideation vignette, crisis-intervention messages appeared in 0/16 responses that included normal objective findings and 16/16 without them
- Across 14 suicidal-ideation vignettes the 988 'Help is available' interstitial fired in 4; the other 10 produced no safety alert in any of 160 responses, and it fired more reliably when no means of self-harm was identified
Methodology Notes
Responses collected via the ChatGPT Health web interface (gpt-5-mini thinking backbone) on 9-11 January 2026, one new thread per condition, no regeneration; 2x2x2x2 within-vignette factorial (anchoring, access barrier, race, sex); gold standard from three physicians; cluster bootstrap and mixed-effects logistic regression with Holm correction. Single standardized prompt template, forced four-level answer (A monitor at home to D emergency department now); no prompt-sensitivity analysis (authors' stated limitation). Synthetic vignettes, no human participants. Data and prompts on Zenodo 10.5281/zenodo.18451491. Received 15 Jan 2026, accepted 20 Feb, published online 23 Feb 2026; print Nat Med 32(5):1671-1675. A published re-analysis (arXiv 2603.11413v4) reports that the headline emergency rate is sensitive to answer format and adjudication. Two later documents bear on the findings and are linked as sources. A Macquarie University re-analysis (arXiv 2603.11413, v4 posted 2026-10-02, not peer reviewed) reports that the emergency under-triage rate depends on the forced four-option answer format and did not reproduce as a stable cross-model finding when the 60 released vignettes were rerun on six API models; it did not test the ChatGPT Health product. A medRxiv preprint from the same Mount Sinai group (10.64898/2026.07.21.26358588, posted 2026-07-23) tested ChatGPT Health on 255 cases, including real emergency-department and nurse-line cases, in single-turn and multi-turn modes and reports that discordant recommendations skewed to lower acuity in both modes (70.8% and 69.0% under-triage).
Sources
Nature Medicine (Brief Communication)(opens in a new tab) (primary)
Zenodo: vignettes, prompts and responses(opens in a new tab) (23 Feb 2026)
Nature Medicine Research Briefing: ChatGPT Health triage advice falls short in key cases(opens in a new tab) (7 May 2026)
Same group, multi-turn follow-up with real cases (medRxiv)(opens in a new tab) (23 Jul 2026)
Re-analysis and replication (Fraile Navarro et al., arXiv v4)(opens in a new tab) (2 Oct 2026)
Mount Sinai follow-up: multi-turn ChatGPT Health triage on real and synthetic cases (medRxiv)(opens in a new tab) (23 Jul 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Ashwin Ramaswamy, Alvira Tyagi, Hannah Hugo, Joy Jiang, Pushkala Jayaraman, Mateen Jangda, Alexis E. Te, Steven A. Kaplan, Joshua Lampert, Robert Freeman, Nicholas Gavin, Ashutosh K. Tewari, Ankit Sakhuja, Bilal Naved, Alexander W. Charney, Mahmud Omar, Michael A. Gorin, Eyal Klang, Girish N. Nadkarni
Tags
Cite This
APA
Ashwin Ramaswamy et al. (2026). ChatGPT Health performance in a structured test of triage recommendations. Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai. https://www.nature.com/articles/s41591-026-04297-7
Related Insights
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
arXiv (Stanford University) · 8 Sept 2026
Single-turn emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026 · 29 Sept 2026