Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
Introduces MedAgent, a framework for synthetically generating realistic multi-turn mental-health sensemaking conversations, the resulting MHSD dataset of more than 2,200 patient-LLM conversations, and MultiSenseEval, a human-centric evaluation framework for multi-turn healthcare conversation. Frontier reasoning models perform below par on patient-centric communication and reach only about 31% average accuracy on precise diagnostic judgements, with performance varying by patient persona and dropping as conversations lengthen.
Publisher
Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Georgia Institute of Technology
Published
1 Jul 2026
Added
today
Key Findings
- The MHSD dataset comprises more than 2,200 synthetic patient-LLM mental-health conversations (2,284 per the paper body: 1,142 each with OpenAI o1 and DeepSeek-R1 as the sensemaking model and GPT-4o acting as the patient, grounded in published case studies).
- Frontier reasoning models yielded below-par patient-centric communication and struggled at precise ('hard') diagnostic capability with average accuracy of about 31%.
- Performance varied with the patient's persona and dropped as the number of conversational turns increased.
- MultiSenseEval evaluates multi-turn conversation ability with human-centric criteria beyond diagnostic accuracy and win-rates; human validation on 100 conversations showed high agreement with the LLM judge.
Methodology Notes
Synthetic conversation generation (LLM patient actor grounded in case studies), evaluation of two frontier reasoning models as sensemakers with a human-centric rubric and LLM judging validated on 100 conversations; no real patients. Published in the ACL 2026 Long Papers volume (July 2026, San Diego), pages 46648 to 46682, DOI 10.18653/v1/2026.acl-long.2164. Verified at the ACL Anthology (bib record and landing-page abstract) on 2026-09-15; the per-model counts and turn-count degradation figures come from the benchmarks beat's read of the PDF.
Sources
Authors
Mohit Chandra, Siddharth Sriraman, Harneet Singh Khanuja, Yiqiao Jin, Munmun De Choudhury
Tags
Cite This
APA
Mohit Chandra et al. (2026). Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations. Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Georgia Institute of Technology. https://aclanthology.org/2026.acl-long.2164/
Related Insights
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
Association for Computational Linguistics (Findings of ACL 2026); Graduate School of Data Science, Seoul National University · 1 Jul 2026
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
arXiv (Texas A&M University; University of Cincinnati) · 8 Sept 2026