MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
MindEval is a framework, designed with PhD-level licensed clinical psychologists, for automatically evaluating language models in realistic multi-turn mental-health therapy conversations. It combines simulated patients with LLM-based evaluation, validates the realism of the simulated patients against human-generated text, and reports strong correlation between automatic and human expert judgments. Twelve state-of-the-art LLMs were evaluated; all scored below 4 out of 6 on average, with the weakest performance on problematic AI-specific communication patterns, and performance deteriorated with longer interactions and with patients presenting severe symptoms. Code, prompts and human evaluation data are released.
Publisher
arXiv (Sword Health AI Research)
Published
23 Nov 2025
Added
today
Key Findings
- All 12 evaluated LLMs average below 4 of 6 on the clinician-designed rubric
- Particular weaknesses in problematic AI-specific patterns of communication (sycophancy, over-validation, reinforcement of maladaptive beliefs)
- Reasoning capability and model scale do not guarantee better performance
- Performance deteriorates as interactions lengthen and when the simulated patient has severe symptoms
- Simulated-patient realism validated against human text; automatic scores correlate strongly with human expert judgments
- Code, prompts and human evaluation data released (github.com/SWORDHealth/mind-eval; default 10-turn interactions)
Methodology Notes
Patient-simulation plus LLM-judge framework co-designed with licensed clinical psychologists; 12 LLMs evaluated; fully automated and model-agnostic by design, so results rest on simulated patients and LLM judging (human evaluation data released for validation). Authored by the vendor Sword Health's research team. arXiv v1 23 Nov 2025, v2 25 Nov 2025, v3 (current) 5 Dec 2025. No version of record found (Crossref title query no match; OpenAlex holds only the arXiv record). Official code at github.com/SWORDHealth/mind-eval (last push 2025-12-05). Verified at the arXiv abstract page (HTTP 200; title, six authors and version history matched); a coverage miss carried on the watchlist since 2026-08-13.
Sources
arXiv abstract page (v3)(opens in a new tab) (primary)
Official code (Sword Health AI Research)(opens in a new tab) (5 Dec 2025)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
José Pombal, Maya D'Eon, Nuno M. Guerreiro, Pedro Henrique Martins, António Farinhas, Ricardo Rei
Tags
Cite This
APA
José Pombal et al. (2025). MindEval: Benchmarking Language Models on Multi-turn Mental Health Support. arXiv (Sword Health AI Research). https://arxiv.org/abs/2511.18491
Related Insights
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut) · 14 Sept 2026
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
arXiv (National University of Singapore, Department of Computer Science) · 3 Sept 2026
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Georgia Institute of Technology · 1 Jul 2026
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
Association for Computational Linguistics (EMNLP 2025) · 1 Nov 2025
Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot
University of Nottingham, School of Computer Science (Responsible AI UK Cornerstone 2 AI Assurance programme) · 18 Sept 2026