HealthBench: Evaluating Large Language Models Towards Improved Human Health
Open-source benchmark from OpenAI that measures the performance and safety of large language models in health conversations. It consists of 5,000 multi-turn conversations between a model and an individual user or a healthcare professional, each graded against conversation-specific rubrics written by 262 physicians, with a model-based grader validated against physician judgment. Results are reported across seven themes, including emergency referrals, context seeking and responding under uncertainty, and five behavioural axes.
Publisher
OpenAI (arXiv preprint)
Published
13 May 2025
Added
today
DOI
—
Key Findings
- 5,000 conversations scored against 48,562 unique rubric criteria; physicians from 26 specialties with practice experience in 60 countries wrote the rubrics.
- Theme distribution: global health 1,097 (21.9%), responding under uncertainty 1,071 (21.4%), expertise-tailored communication 919 (18.4%), context seeking 594 (11.9%), emergency referrals 482 (9.6%), health data tasks 477 (9.5%), response depth 360 (7.2%).
- Overall scores ranged from 0.16 (GPT-3.5 Turbo) to 0.60 (o3); emergency referrals and expertise-tailored communication scored highest, while context seeking, health data tasks and global health lagged.
- HealthBench Consensus covers 34 physician-consensus criteria (8,053 occurrences); on HealthBench Hard no evaluated model scored above 32%.
- Models outperformed an unassisted physician baseline, and model-physician grading agreement was similar to physician-physician agreement.
Methodology Notes
Most conversations were synthetically generated from physician-written prompt seeds; another subset comes from physician red-teaming of language models in health settings. A model-based grader scores each response criterion by criterion; the authors report meta-evaluation against physician grading. Models evaluated: GPT-3.5 Turbo, GPT-4o, GPT-4.1, o1, o3, Claude 3.7 Sonnet, Gemini 2.5 Pro (Mar 2025), Grok 3 and Llama 4 Maverick. Data and code are released through OpenAI's simple-evals repository (per the paper). Developer-authored benchmark in which the developer's own model scores highest; synthetic conversations may not reflect real user traffic. arXiv v1 submitted 2025-05-13; no later version.
Sources
arXiv preprint(opens in a new tab) (primary)
Full paper (PDF, v1)(opens in a new tab) (13 May 2025)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, Karan Singhal
Tags
Cite This
APA
Rahul K. Arora et al. (2025). HealthBench: Evaluating Large Language Models Towards Improved Human Health. OpenAI (arXiv preprint). https://arxiv.org/abs/2505.08775
Related Insights
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026
Making medical AI benchmarks clinically interpretable: the case of mental health
BMJ Mental Health (BMJ); RAND · 28 Sept 2026
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
OpenAI · 23 Sept 2026
GPT-6 Astra System Card
OpenAI · 3 Sept 2026