Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified data and each paper's default judges and simulated users. The leaderboard has a quality table (MentalHealthBench, MindEval, HealthBench-Psych, EQ-Bench 3, CounselBench EVAL and Adv, CBT-Bench) and a safety table (VERA-MH, Spiral-Bench, SIM-VAIL). Anyone can submit results as pull requests. The first results, for nine models, were submitted by the maintainer between 1 and 3 October 2026.

Publisher

Slingshot AI

Published

1 Oct 2026

Added

today

DOI

—

Key Findings

  • Safety table, nine models: GPT-6 Astra ranks first (VERA-MH 81.2, Spiral-Bench 80.5, SIM-VAIL 1.027 where lower is better) and Gemini 2.5 Pro last (30.7, 44.6, 3.599)
  • Share of VERA-MH suicidal-ideation conversations (200 per model, up to 30 turns) with at least one high-harm rating: GPT-6 Luna 6.95%, GPT-6 Astra 10.61%, GPT-5 21.51%, Claude Opus 5.5 21.55%, Claude Sonnet 4.5 55.8%, Qwen3.5-122B-A10B 73.12%, Qwen3.5-27B 77.96%, Gemini 2.5 Pro 90.56%, GPT-4o 97.18%
  • GPT-6 Astra's lowest VERA-MH dimension is Follows AI Boundaries (dimension score 44.45; 2 of 176 judgments best practice, 11 damaging), against 100 on Confirms Risk and 94.97 on Detects Potential Risk
  • Spiral-Bench records the highest sycophancy value for Gemini 2.5 Pro (3.288) and the highest delusion-reinforcement value for Qwen3.5-27B (2.152); GPT-6 Luna records 0.0 on both
  • Quality table win rates: Claude Opus 5.5 86%, GPT-6 Astra 77%, GPT-6 Luna 64%, GPT-4o 20%

Methodology Notes

Software harness (MIT licence; v0.1.0 released 2026-10-01, v0.1.1 2026-10-02) plus a static site built from per-model result files in github.com/slingshot-ai/mheval-leaderboard (repository created 2026-10-01; site reads 'Updated October 3, 2026'). Each benchmark keeps its original judge and simulated users. According to the harness README, VERA-MH uses the recommended SI profile with a gpt-5.4 judge and gpt-5.2/claude-opus-4-5 simulated users, Spiral-Bench uses a gpt-5 judge and SIM-VAIL a claude-opus-4.5 judge, so same-family judge preference may affect rankings. Gemini 2.5 Pro, GPT-4o and Claude Sonnet 4.5 are older releases than the other entries. The maintainer self-reported the results (the GPT-6 Astra VERA-MH file is dated 2026-10-02). There is no paper and no peer review. The leaderboard is live and changes as submissions merge. Figures were read on 2026-10-05 from the rendered site and from results/<model>/vera_mh.json and spiral_bench.json.

Authors

Ziyi Zhu

Tags

leaderboardmhevalslingshot-aivera-mhspiral-benchsim-vailgpt-6-astraclaude-opus-5-5

Cite This

APA

Ziyi Zhu. (2026). Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard. Slingshot AI. https://slingshot-ai.github.io/mheval-leaderboard/