From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
DischargeBench, a persona-grounded simulation in which a candidate language model conducts a multi-turn discharge-education session with a virtual patient while a monitor agent keeps the patient realistic without touching the educator. The benchmark (MIMIC-IV-Ext-DischargeBench) has 477 cases across 24 ICD chapters with persona axes for personality, education level, health literacy and recall of past medical history, and scores conversation quality, topic checklist, patient comprehension and factual consistency with a language-model judge aligned to physician annotations.
Publisher
arXiv (University of Massachusetts Lowell; University of Massachusetts Amherst)
Published
22 Jul 2026
Added
today
DOI
—
Key Findings
- 477 MIMIC-IV cases over 24 ICD chapters with four persona axes; four scoring axes judged by a language model aligned against annotations from two licensed emergency physicians
- Aggregate scores across closed- and open-source models conceal clinically relevant variation across ICD chapters and patient personas
- Difficult personas expose coverage failures, comprehension gaps and reduced source-answer agreement
- The authors argue evaluation of discharge education should centre patient understanding rather than text quality or answer accuracy
Methodology Notes
Simulation-based benchmark with model-simulated patients and a model judge; physician alignment on a subset of cases. Submitted 22 July 2026, announced 21 September 2026 (date is the arXiv v1 submission date). CC BY. The underlying notes require MIMIC-IV credentialed access.
Sources
arXiv preprint(opens in a new tab) (primary)
Authors
Jang, Won Seok, Yao, Zonghai, Yu, Hong
Tags
Cite This
APA
Jang, Won Seok, Yao, Zonghai, Yu, Hong. (2026). From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators. arXiv (University of Massachusetts Lowell; University of Massachusetts Amherst). https://arxiv.org/abs/2609.20827
Related Insights
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
arXiv (Stanford University) · 8 Sept 2026
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California) · 15 Apr 2025
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University · 14 Jul 2026