Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios

A factorial safety evaluation of consumer chatbot advice in a time-critical surgical emergency. Thirty patient-style questions about suspected acute appendicitis were put to three public assistants under two prompting conditions with three repeats, producing 540 responses, each coded by three blinded surgeons for major harm, meaning advice plausibly producing delayed diagnosis, inappropriate treatment or significant harm. Harm rates were non-trivial in every question category, structured prompting did not consistently reduce them, and the three products behaved similarly.

Publisher

BMC Medical Informatics and Decision Making; Nigde Omer Halisdemir University; Ankara Bilkent City Hospital; Kutahya Dumlupinar University

Published

29 Aug 2026

Added

today

Key Findings

  • 540 responses from a predefined factorial design: 30 patient-style questions x 3 AI systems x 2 prompting conditions x 3 repeats.
  • Medication questions produced harmful advice most often: 28.8% under the any-evaluator definition and 23.8% under the majority-evaluator definition.
  • Symptom interpretation 19.4% / 16.8%; the safety of waiting at home 15.9% / 13.5%.
  • Structured prompting did not consistently reduce major harm, so prompt engineering is not a mitigation for this failure.
  • Safety profiles were broadly similar across ChatGPT (GPT-5.2 Instant), Gemini 3 Flash and Microsoft Copilot, so there is no safer product to route users to.
  • Worst-case repeat analysis shows harmful advice recurring across repeated responses under identical conditions rather than appearing as a sampling fluke; comorbidity and pregnancy scenarios were most likely to draw majority-rated major harm.

Methodology Notes

Three blinded surgeons coded each response against a prespecified major-harm definition, reported under both any-evaluator and majority-evaluator rules. The authors state no formal power-based sample-size calculation was performed. One clinical condition, one language, and a Turkish surgical team's reference standard, so the absolute rates should not be generalised beyond appendicitis. Queries were run 2025-12-18 to 2026-01-29. Published 2026-08-29, CC BY 4.0, DOI 10.1186/s12911-026-03800-x; the article page was fetched and the abstract read.

Authors

Ali Riza Erdogan, Fahri Martli, Mehmet Fatih Ekici, Pinar Erdogan

Tags

appendicitisemergencypatient-facingprompt-engineeringturkeyharm-rate

Cite This

APA

Ali Riza Erdogan et al. (2026). Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios. BMC Medical Informatics and Decision Making; Nigde Omer Halisdemir University; Ankara Bilkent City Hospital; Kutahya Dumlupinar University. https://doi.org/10.1186/s12911-026-03800-x