Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models
Cross-sectional evaluation of six consumer chatbots answering 56 public-consultation questions about sudden cardiac death, with five blinded senior cardiologists rating each of the 336 responses for safety, accuracy and empathy, plus DISCERN, EQIP, JAMA and Global Quality Score instruments and readability formulas. Every chatbot produced some responses with potential safety concerns, most often delayed advice to activate emergency services for chest pain, syncope or post-viral symptoms, and over-reassurance about prevention.
Publisher
BMC Public Health (BMC, Springer Nature); Bozhou People's Hospital
Published
6 Oct 2026
Added
today
Key Findings
- All six chatbots produced responses containing potential safety concerns; ChatGPT Plus 5.5 and Perplexity had the highest proportions of safe responses and Doubao the lowest, and most flagged responses carried a single concern.
- Recurring safety problems were delayed emergency activation for chest pain, syncope or post-viral symptoms, and over-reassurance about prevention.
- Inter-rater agreement was high for safety (Fleiss' kappa 0.856) with intraclass correlation coefficients of 0.801 to 0.887 across dimensions; accuracy differed across models (Kendall's W 0.864) with Doubao lower, and empathy differed (Kendall's W 0.531) with DeepSeek-V4-pro scoring highest.
- DeepSeek-V4-pro scored highest on DISCERN and EQIP and had the most favourable readability profile, while all models required relatively high reading grades; Microsoft Copilot scored highest on the JAMA criteria.
- The authors conclude the tools may support general education about sudden cardiac death but should not replace emergency medical services or clinician assessment.
Methodology Notes
56 standardised questions submitted once each in English to ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao and Perplexity through their public interfaces from China between 2026-05-10 and 2026-05-16 (336 response items); five blinded senior cardiology raters. The authors describe the results as a time-, language- and access-specific snapshot rather than a stable ranking. Received 2026-06-17, accepted 2026-09-30, published 2026-10-06; CC BY-NC-ND 4.0. Verification: Crossref record and the full text on link.springer.com (read through the publisher's cookie-redirect chain; bmcpublichealth.biomedcentral.com redirects there); no Europe PMC record yet.
Sources
BMC Public Health article (DOI)(opens in a new tab) (primary)
Full text on SpringerLink(opens in a new tab) (6 Oct 2026)
Authors
Youyou Chen, Hui Ma, Huimin Wang, Duoxue Chen, Haiyan Wang, Rongyan Jiang
Tags
Cite This
APA
Youyou Chen et al. (2026). Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models. BMC Public Health (BMC, Springer Nature); Bozhou People's Hospital. https://doi.org/10.1186/s12889-026-29766-z
Related Insights
ChatGPT Health performance in a structured test of triage recommendations
Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai · 23 Feb 2026
HealthBench: Evaluating Large Language Models Towards Improved Human Health
OpenAI (arXiv preprint) · 13 May 2025