MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
A multi-turn diagnostic benchmark built from 1,387 board-exam cases across 17 specialties, each converted to a 24-slot clinical record and then played out as doctor-patient dual-agent dialogues under varying patient personas with the clinical content held fixed, with cases classified by difficulty using model-based uncertainty. Across fifteen LLM doctor agents, moving from single-turn diagnosis on the standardized record to a multi-turn consultation costs roughly 13 to 39 accuracy points, more turns improve question relevance but accuracy plateaus after 6 to 12 turns, and patient persona alone shifts accuracy by about 7 to 8 points from lowest to highest education level.
Publisher
arXiv (MBZUAI; Cairo University; Ain Shams University; CSIRO; INSAIT, Sofia University)
Published
11 Sept 2026
Added
today
Key Findings
- Moving from single-turn diagnosis on a standardized record to a multi-turn consultation caused accuracy degradations of roughly 13 to 39 points across fifteen LLM doctor agents.
- More turns reliably improved question relevance, but diagnostic accuracy showed diminishing returns and typically plateaued after 6 to 12 turns.
- Patient persona differences shifted diagnostic accuracy by about 7 to 8 points between the lowest and highest patient education levels, with the clinical content held fixed, which the authors flag as an equity risk single-turn benchmarks miss.
- The benchmark covers 1,387 board-exam cases in 17 specialties, each expanded into a structured 24-slot clinical record and instantiated as controlled dual-agent dialogues.
Methodology Notes
Simulated dual-agent doctor-patient dialogues with persona variation and model-uncertainty-based difficulty stratification; fifteen LLM doctor agents (not named in the abstract); no real patients. arXiv v1 2026-09-11 (14 September partition). Affiliations from the PDF title block. Overlaps in aim with VeriSim (patient-communication noise), which is held separately; the two are cross-referenced rather than merged.
Sources
Authors
Youssef Mohamed, Ahmed Heakl, Qinrong Cui, Junhong Liang, Rafiq Ali, Bdour Babillie, Nazira Dunbayeva, Lang Gao, Omar Hussein, Ahmed Nada, Ahmed Mohamed Magdy Mohamed, Jinghui Liu, Salman Khan, Imran Razzak, Yuxia Wang, Xiuying Chen
Tags
Cite This
APA
Youssef Mohamed et al. (2026). MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations. arXiv (MBZUAI; Cairo University; Ain Shams University; CSIRO; INSAIT, Sofia University). https://arxiv.org/abs/2609.12851
Related Insights
VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise
arXiv (George Mason University); accepted to Findings of EMNLP 2026 · 12 Apr 2026
Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency
arXiv (single author, no institutional affiliation stated) · 2 Jun 2026