Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Peer-reviewed MLHC 2026 paper (version of record of the July 2026 arXiv preprint) testing whether large language models identify and correct false presuppositions in patients' questions as a conversation continues. It introduces ThReadMed-QA, 2,437 patient-physician conversation threads (8,204 question-answer pairs) derived from real patient interactions on r/AskDocs, and scores five LLMs with a rubric-based LLM-as-a-judge. Frontier models correct misconceptions on most initial questions but degrade substantially over subsequent turns, and an oracle analysis attributes much of the drop to propagation of the model's own earlier errors.

Publisher

Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University

Published

8 Sept 2026

Added

today

DOI

Key Findings

  • Even frontier models that address misconceptions in a single interaction degrade substantially over subsequent turns; the authors report correction on roughly 85% of initial questions falling to roughly 50% within two follow-ups for GPT-5 and Claude Haiku
  • Replacing prior model outputs with physician responses in an oracle condition recovers much of the loss, indicating that error propagation drives most of the degradation, while performance stays imperfect even with correct context
  • ThReadMed-QA comprises 2,437 real patient-physician threads (8,204 question-answer pairs) from r/AskDocs with verified physician answers; the Hugging Face release is gated

Methodology Notes

Multi-turn evaluation of five LLMs on ThReadMed-QA with a rubric-based LLM judge; an oracle condition substitutes physician responses for prior model turns. Companion dataset paper by the same authors: arXiv 2603.11281 (2026-03-11). Version of record in PMLR volume 340 (MLHC 2026, conference held at Johns Hopkins University School of Medicine; volume published 8 September 2026); PMLR does not register DOIs promptly, so no DOI yet. Verified by fetching the PMLR abstract page (HTTP 200; title and both authors in the citation metadata) and the volume index. Supersedes the held arXiv preprint row (2607.12884, v1 2026-07-14).

Authors

Monica Munnangi, Saiph Savage

Tags

mlhc-2026threadmed-qaaskdocsmulti-turnmisconceptionsllm-as-judgeversion-of-record

Cite This

APA

Monica Munnangi, Saiph Savage. (2026). Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations. Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University. https://proceedings.mlr.press/v340/munnangi26a.html