Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Peer-reviewed MLHC 2026 paper (version of record of the July 2026 arXiv preprint) testing whether large language models identify and correct false presuppositions in patients' questions as a conversation continues. It introduces ThReadMed-QA, 2,437 patient-physician conversation threads (8,204 question-answer pairs) derived from real patient interactions on r/AskDocs, and scores five LLMs with a rubric-based LLM-as-a-judge. Frontier models correct misconceptions on most initial questions but degrade substantially over subsequent turns, and an oracle analysis attributes much of the drop to propagation of the model's own earlier errors.
Publisher
Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University
Published
8 Sept 2026
Added
today
DOI
—
Key Findings
- Even frontier models that address misconceptions in a single interaction degrade substantially over subsequent turns; the authors report correction on roughly 85% of initial questions falling to roughly 50% within two follow-ups for GPT-5 and Claude Haiku
- Replacing prior model outputs with physician responses in an oracle condition recovers much of the loss, indicating that error propagation drives most of the degradation, while performance stays imperfect even with correct context
- ThReadMed-QA comprises 2,437 real patient-physician threads (8,204 question-answer pairs) from r/AskDocs with verified physician answers; the Hugging Face release is gated
Methodology Notes
Multi-turn evaluation of five LLMs on ThReadMed-QA with a rubric-based LLM judge; an oracle condition substitutes physician responses for prior model turns. Companion dataset paper by the same authors: arXiv 2603.11281 (2026-03-11). Version of record in PMLR volume 340 (MLHC 2026, conference held at Johns Hopkins University School of Medicine; volume published 8 September 2026); PMLR does not register DOIs promptly, so no DOI yet. Verified by fetching the PMLR abstract page (HTTP 200; title and both authors in the citation metadata) and the volume index. Supersedes the held arXiv preprint row (2607.12884, v1 2026-07-14).
Authors
Monica Munnangi, Saiph Savage
Tags
Cite This
APA
Monica Munnangi, Saiph Savage. (2026). Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations. Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University. https://proceedings.mlr.press/v340/munnangi26a.html
Related Insights
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University · 14 Jul 2026
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
Association for Computational Linguistics (Findings of ACL 2026); Duke University; Stanford University · 1 Jul 2026
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
arXiv (MBZUAI; Cairo University; Ain Shams University; CSIRO; INSAIT, Sofia University) · 11 Sept 2026
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California) · 15 Apr 2025
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Verily Health · 8 Sept 2026