Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified data and each paper's default judges and simulated users. The leaderboard has a quality table (MentalHealthBench, MindEval, HealthBench-Psych, EQ-Bench 3, CounselBench EVAL and Adv, CBT-Bench) and a safety table (VERA-MH, Spiral-Bench, SIM-VAIL). Anyone can submit results as pull requests. The first results, for nine models, were submitted by the maintainer between 1 and 3 October 2026.
Publisher
Slingshot AI
Published
1 Oct 2026
Added
today
DOI
—
Key Findings
- Safety table, nine models: GPT-6 Astra ranks first (VERA-MH 81.2, Spiral-Bench 80.5, SIM-VAIL 1.027 where lower is better) and Gemini 2.5 Pro last (30.7, 44.6, 3.599)
- Share of VERA-MH suicidal-ideation conversations (200 per model, up to 30 turns) with at least one high-harm rating: GPT-6 Luna 6.95%, GPT-6 Astra 10.61%, GPT-5 21.51%, Claude Opus 5.5 21.55%, Claude Sonnet 4.5 55.8%, Qwen3.5-122B-A10B 73.12%, Qwen3.5-27B 77.96%, Gemini 2.5 Pro 90.56%, GPT-4o 97.18%
- GPT-6 Astra's lowest VERA-MH dimension is Follows AI Boundaries (dimension score 44.45; 2 of 176 judgments best practice, 11 damaging), against 100 on Confirms Risk and 94.97 on Detects Potential Risk
- Spiral-Bench records the highest sycophancy value for Gemini 2.5 Pro (3.288) and the highest delusion-reinforcement value for Qwen3.5-27B (2.152); GPT-6 Luna records 0.0 on both
- Quality table win rates: Claude Opus 5.5 86%, GPT-6 Astra 77%, GPT-6 Luna 64%, GPT-4o 20%
Methodology Notes
Software harness (MIT licence; v0.1.0 released 2026-10-01, v0.1.1 2026-10-02) plus a static site built from per-model result files in github.com/slingshot-ai/mheval-leaderboard (repository created 2026-10-01; site reads 'Updated October 3, 2026'). Each benchmark keeps its original judge and simulated users. According to the harness README, VERA-MH uses the recommended SI profile with a gpt-5.4 judge and gpt-5.2/claude-opus-4-5 simulated users, Spiral-Bench uses a gpt-5 judge and SIM-VAIL a claude-opus-4.5 judge, so same-family judge preference may affect rankings. Gemini 2.5 Pro, GPT-4o and Claude Sonnet 4.5 are older releases than the other entries. The maintainer self-reported the results (the GPT-6 Astra VERA-MH file is dated 2026-10-02). There is no paper and no peer review. The leaderboard is live and changes as submissions merge. Figures were read on 2026-10-05 from the rendered site and from results/<model>/vera_mh.json and spiral_bench.json.
Sources
Mental Health Evaluation Leaderboard (Slingshot AI)(opens in a new tab) (primary)
mheval harness repository (README, benchmark table)(opens in a new tab) (1 Oct 2026)
mheval v0.1.0 release(opens in a new tab) (1 Oct 2026)
Leaderboard repository with per-model result files(opens in a new tab) (3 Oct 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Ziyi Zhu
Tags
Cite This
APA
Ziyi Zhu. (2026). Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard. Slingshot AI. https://slingshot-ai.github.io/mheval-leaderboard/
Related Insights
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
OpenAI · 23 Sept 2026
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
arXiv (Sword Health AI Research) · 23 Nov 2025
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut) · 14 Sept 2026