Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study

Prospective, conversation-level evaluation of jAImes, a retrieval-augmented Parkinson's disease information chatbot commissioned by Parkinson Stiftung Deutschland and deployed publicly in Germany, across its first 129 days of operation (11 November 2025 to 20 March 2026). The authors analyse all 2,035 conversations (6,146 messages) and introduce CARE-LLM (Conversation-level AI Real-world Evaluation), a four-component post-market surveillance framework combining automated triage of every conversation, structured expert review of flagged cases, sampling-based sensitivity validation against the conversations rated good, and a failure-class feedback loop. Favourable automated classifications coexisted with clinically critical failures that only expert conversation-level review detected.

Publisher

The Lancet Regional Health – Europe (Elsevier); University Hospital Würzburg, Department of Neurology; Parkinson Stiftung Deutschland

Published

12 Sept 2026

Added

today

Key Findings

  • AI-assisted triage rated 1,803 of 2,035 conversations (88.6%) good, 224 (11.0%) partially adequate and 8 (0.4%) inadequate
  • Expert review of 45 flagged conversations confirmed five critical events; independent re-review of 100 randomly sampled conversations rated good found four (95% CI 1.6 to 9.8) with clinically critical errors the triage had missed
  • Confirmed critical events spanned three failure classes (knowledge boundary, robustness, escalation) and included an inappropriate memantine recommendation in Parkinson's disease dementia and explicit suicidal ideation that did not trigger the intended emergency response
  • The system was scoped as non-diagnostic and non-therapeutic, answered only from a curated knowledge base, carried a safety prompt prohibiting dosing and therapy changes, and had a predefined emergency-response pathway; the failures occurred inside that design
  • The authors report finding no prior prospective conversation-level safety-monitoring study of an autonomously deployed patient-facing LLM system in any disease area

Methodology Notes

Single deployed system, single country, unrestricted public use; all conversations analysed; critical-event adjudication described as structurally independent of the developer and the commissioning foundation. Declared interests: one author is founder and managing director of smardis.tech, the paid technical developer of jAImes; the senior author is president of Parkinson Stiftung Deutschland, which commissioned and operates the system. Funding: DFG (TRR 295), IZKF Würzburg, BMBF; ethics review University of Würzburg. Received 15 July 2026, accepted 2 September 2026, published online 12 September 2026; CC BY 4.0. The Lancet site blocks automated access; the full text was read from the Europe PMC deposit PMC13587714 and the abstract confirmed on PubMed 42761851.

Authors

Lange, Florian, Mardi, Ssaman, Binder, Tobias, Reich, Martin M., Odorfer, Thorsten, Volkmann, Jens

Tags

care-llmjaimesparkinsonspost-market-surveillanceraggermanyescalation-failure

Cite This

APA

Lange, Florian et al. (2026). Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study. The Lancet Regional Health – Europe (Elsevier); University Hospital Würzburg, Department of Neurology; Parkinson Stiftung Deutschland. https://doi.org/10.1016/j.lanepe.2026.101866