Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Preprint, marked accepted to MLHC 2026, testing whether large language models identify and correct false presuppositions in patients' questions as a conversation continues. It introduces ThReadMed-QA, 2,437 patient-physician conversation threads (8,204 question-answer pairs) derived from real patient interactions on r/AskDocs, and scores five LLMs with a rubric-based LLM-as-a-judge. Frontier models correct misconceptions on most initial questions but degrade substantially over subsequent turns, and an oracle analysis attributes much of the drop to propagation of the model's own earlier errors.

Publisher

arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University

Published

14 Jul 2026

Added

today

Key Findings

  • Even frontier models that address misconceptions in a single interaction degrade substantially over subsequent turns; the authors report correction on roughly 85% of initial questions falling to roughly 50% within two follow-ups for GPT-5 and Claude Haiku
  • Replacing prior model outputs with physician responses in an oracle condition recovers much of the loss, indicating error propagation drives most of the degradation, while performance stays imperfect even with correct context
  • ThReadMed-QA comprises 2,437 real patient-physician threads (8,204 QA pairs) from r/AskDocs with verified physician answers; the Hugging Face release is gated

Methodology Notes

Multi-turn evaluation of five LLMs on ThReadMed-QA with a rubric-based LLM judge; oracle condition substitutes physician responses for prior model turns. Companion dataset paper by the same authors: arXiv 2603.11281 (2026-03-11). arXiv v1 2026-07-14; comment 'Accepted to MLHC 2026'; no PMLR volume yet. Curator fetched the abstract page on 2026-09-17; the beat confirmed Northeastern affiliations via OpenAlex.

Authors

Munnangi, Monica, Savage, Saiph

Tags

mlhc-2026threadmed-qaaskdocsmulti-turnmisconceptionsllm-as-judge

Cite This

APA

Munnangi, Monica, Savage, Saiph. (2026). Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations. arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University. https://arxiv.org/abs/2607.12884