Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

MindEval: Benchmarking Language Models on Multi-turn Mental Health Support

MindEval is a framework, designed with PhD-level licensed clinical psychologists, for automatically evaluating language models in realistic multi-turn mental-health therapy conversations. It combines simulated patients with LLM-based evaluation, validates the realism of the simulated patients against human-generated text, and reports strong correlation between automatic and human expert judgments. Twelve state-of-the-art LLMs were evaluated; all scored below 4 out of 6 on average, with the weakest performance on problematic AI-specific communication patterns, and performance deteriorated with longer interactions and with patients presenting severe symptoms. Code, prompts and human evaluation data are released.

Publisher

arXiv (Sword Health AI Research)

Published

23 Nov 2025

Added

today

Key Findings

  • All 12 evaluated LLMs average below 4 of 6 on the clinician-designed rubric
  • Particular weaknesses in problematic AI-specific patterns of communication (sycophancy, over-validation, reinforcement of maladaptive beliefs)
  • Reasoning capability and model scale do not guarantee better performance
  • Performance deteriorates as interactions lengthen and when the simulated patient has severe symptoms
  • Simulated-patient realism validated against human text; automatic scores correlate strongly with human expert judgments
  • Code, prompts and human evaluation data released (github.com/SWORDHealth/mind-eval; default 10-turn interactions)

Methodology Notes

Patient-simulation plus LLM-judge framework co-designed with licensed clinical psychologists; 12 LLMs evaluated; fully automated and model-agnostic by design, so results rest on simulated patients and LLM judging (human evaluation data released for validation). Authored by the vendor Sword Health's research team. arXiv v1 23 Nov 2025, v2 25 Nov 2025, v3 (current) 5 Dec 2025. No version of record found (Crossref title query no match; OpenAlex holds only the arXiv record). Official code at github.com/SWORDHealth/mind-eval (last push 2025-12-05). Verified at the arXiv abstract page (HTTP 200; title, six authors and version history matched); a coverage miss carried on the watchlist since 2026-08-13.

Authors

José Pombal, Maya D'Eon, Nuno M. Guerreiro, Pedro Henrique Martins, António Farinhas, Ricardo Rei

Tags

mindevalsword-healthpatient-simulationllm-judgemulti-turncoverage-miss

Cite This

APA

José Pombal et al. (2025). MindEval: Benchmarking Language Models on Multi-turn Mental Health Support. arXiv (Sword Health AI Research). https://arxiv.org/abs/2511.18491