Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios
A factorial safety evaluation of consumer chatbot advice in a time-critical surgical emergency. Thirty patient-style questions about suspected acute appendicitis were put to three public assistants under two prompting conditions with three repeats, producing 540 responses, each coded by three blinded surgeons for major harm, meaning advice plausibly producing delayed diagnosis, inappropriate treatment or significant harm. Harm rates were non-trivial in every question category, structured prompting did not consistently reduce them, and the three products behaved similarly.
Publisher
BMC Medical Informatics and Decision Making; Nigde Omer Halisdemir University; Ankara Bilkent City Hospital; Kutahya Dumlupinar University
Published
29 Aug 2026
Added
today
Key Findings
- 540 responses from a predefined factorial design: 30 patient-style questions x 3 AI systems x 2 prompting conditions x 3 repeats.
- Medication questions produced harmful advice most often: 28.8% under the any-evaluator definition and 23.8% under the majority-evaluator definition.
- Symptom interpretation 19.4% / 16.8%; the safety of waiting at home 15.9% / 13.5%.
- Structured prompting did not consistently reduce major harm, so prompt engineering is not a mitigation for this failure.
- Safety profiles were broadly similar across ChatGPT (GPT-5.2 Instant), Gemini 3 Flash and Microsoft Copilot, so there is no safer product to route users to.
- Worst-case repeat analysis shows harmful advice recurring across repeated responses under identical conditions rather than appearing as a sampling fluke; comorbidity and pregnancy scenarios were most likely to draw majority-rated major harm.
Methodology Notes
Three blinded surgeons coded each response against a prespecified major-harm definition, reported under both any-evaluator and majority-evaluator rules. The authors state no formal power-based sample-size calculation was performed. One clinical condition, one language, and a Turkish surgical team's reference standard, so the absolute rates should not be generalised beyond appendicitis. Queries were run 2025-12-18 to 2026-01-29. Published 2026-08-29, CC BY 4.0, DOI 10.1186/s12911-026-03800-x; the article page was fetched and the abstract read.
Authors
Ali Riza Erdogan, Fahri Martli, Mehmet Fatih Ekici, Pinar Erdogan
Tags
Cite This
APA
Ali Riza Erdogan et al. (2026). Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios. BMC Medical Informatics and Decision Making; Nigde Omer Halisdemir University; Ankara Bilkent City Hospital; Kutahya Dumlupinar University. https://doi.org/10.1186/s12911-026-03800-x