Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Preprint, marked accepted to MLHC 2026, testing whether large language models identify and correct false presuppositions in patients' questions as a conversation continues. It introduces ThReadMed-QA, 2,437 patient-physician conversation threads (8,204 question-answer pairs) derived from real patient interactions on r/AskDocs, and scores five LLMs with a rubric-based LLM-as-a-judge. Frontier models correct misconceptions on most initial questions but degrade substantially over subsequent turns, and an oracle analysis attributes much of the drop to propagation of the model's own earlier errors.
Publisher
arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University
Published
14 Jul 2026
Added
today
Key Findings
- Even frontier models that address misconceptions in a single interaction degrade substantially over subsequent turns; the authors report correction on roughly 85% of initial questions falling to roughly 50% within two follow-ups for GPT-5 and Claude Haiku
- Replacing prior model outputs with physician responses in an oracle condition recovers much of the loss, indicating error propagation drives most of the degradation, while performance stays imperfect even with correct context
- ThReadMed-QA comprises 2,437 real patient-physician threads (8,204 QA pairs) from r/AskDocs with verified physician answers; the Hugging Face release is gated
Methodology Notes
Multi-turn evaluation of five LLMs on ThReadMed-QA with a rubric-based LLM judge; oracle condition substitutes physician responses for prior model turns. Companion dataset paper by the same authors: arXiv 2603.11281 (2026-03-11). arXiv v1 2026-07-14; comment 'Accepted to MLHC 2026'; no PMLR volume yet. Curator fetched the abstract page on 2026-09-17; the beat confirmed Northeastern affiliations via OpenAlex.
Authors
Munnangi, Monica, Savage, Saiph
Tags
Cite This
APA
Munnangi, Monica, Savage, Saiph. (2026). Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations. arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University. https://arxiv.org/abs/2607.12884
Related Insights
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
Association for Computational Linguistics (Findings of ACL 2026); Duke University; Stanford University · 1 Jul 2026
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
arXiv (MBZUAI; Cairo University; Ain Shams University; CSIRO; INSAIT, Sofia University) · 11 Sept 2026
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California) · 15 Apr 2025