Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

HealthBench: Evaluating Large Language Models Towards Improved Human Health

Open-source benchmark from OpenAI that measures the performance and safety of large language models in health conversations. It consists of 5,000 multi-turn conversations between a model and an individual user or a healthcare professional, each graded against conversation-specific rubrics written by 262 physicians, with a model-based grader validated against physician judgment. Results are reported across seven themes, including emergency referrals, context seeking and responding under uncertainty, and five behavioural axes.

Publisher

OpenAI (arXiv preprint)

Published

13 May 2025

Added

today

DOI

—

Key Findings

  • 5,000 conversations scored against 48,562 unique rubric criteria; physicians from 26 specialties with practice experience in 60 countries wrote the rubrics.
  • Theme distribution: global health 1,097 (21.9%), responding under uncertainty 1,071 (21.4%), expertise-tailored communication 919 (18.4%), context seeking 594 (11.9%), emergency referrals 482 (9.6%), health data tasks 477 (9.5%), response depth 360 (7.2%).
  • Overall scores ranged from 0.16 (GPT-3.5 Turbo) to 0.60 (o3); emergency referrals and expertise-tailored communication scored highest, while context seeking, health data tasks and global health lagged.
  • HealthBench Consensus covers 34 physician-consensus criteria (8,053 occurrences); on HealthBench Hard no evaluated model scored above 32%.
  • Models outperformed an unassisted physician baseline, and model-physician grading agreement was similar to physician-physician agreement.

Methodology Notes

Most conversations were synthetically generated from physician-written prompt seeds; another subset comes from physician red-teaming of language models in health settings. A model-based grader scores each response criterion by criterion; the authors report meta-evaluation against physician grading. Models evaluated: GPT-3.5 Turbo, GPT-4o, GPT-4.1, o1, o3, Claude 3.7 Sonnet, Gemini 2.5 Pro (Mar 2025), Grok 3 and Llama 4 Maverick. Data and code are released through OpenAI's simple-evals repository (per the paper). Developer-authored benchmark in which the developer's own model scores highest; synthetic conversations may not reflect real user traffic. arXiv v1 submitted 2025-05-13; no later version.

Authors

Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, Karan Singhal

Tags

healthbenchopenaihealth-conversationsrubric-evaluationphysician-rubricsemergency-referrals

Cite This

APA

Rahul K. Arora et al. (2025). HealthBench: Evaluating Large Language Models Towards Improved Human Health. OpenAI (arXiv preprint). https://arxiv.org/abs/2505.08775