mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0)
A clinician-developed benchmark for evaluating how language models behave in simulated suicide-related conversations. Fifty licensed or supervised mental-health clinicians, drawn from a pool of 7,947, authored 300 multi-turn roleplays stratified across four C-SSRS-informed risk levels (no evident risk, low, moderate, high) and conversed with six models in default API configuration without system prompts. Ten clinician evaluators who met a Krippendorff's alpha threshold of 0.80 then labelled every model utterance under a multi-label framework of helpful, less harmful and more harmful behaviours. A severity-weighted 0-10 mPACT-S score separates the models by about 2.4 points, and the authors report that models were better at avoiding harm than at giving actively helpful, clinically appropriate responses.
Publisher
PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington)
Published
15 May 2026
Added
today
Key Findings
- Severity-weighted mPACT-S scores (0-10, higher is better): Claude Sonnet 4.5 9.19, GPT-5.2 9.09, Gemini 2.5 Flash 8.47, Llama 3.3 70B 7.62, Grok 4.1 6.79, Mistral Medium 3 6.79
- Share of conversations containing at least one 'more harmful' utterance: Mistral Medium 3 47%, Grok 4.1 38%, Llama 3.3 70B 29%, Gemini 2.5 Flash 19%, GPT-5.2 16%, Claude Sonnet 4.5 6%
- Claude Sonnet 4.5 scored highest in high-risk conversations (10) and lowest in no-risk ones (8.1), while a Simple Harm Avoidance score ranks GPT-5.2 first, showing that harm avoidance and clinically appropriate helping diverge across systems
- 300 roleplays: 72 no-risk, 73 low, 76 moderate, 79 high; authors averaged 6.6 roleplays each; 104 roleplays (30%) were reviewed for protocol fidelity; evaluators annotated about 25 conversations each
- The benchmark is positioned against LLM-judge and automated-user designs and names VERA-MH and Throughline as comparators; all generation and scoring is human
- The roleplays and rubric are not released; models were selected in March 2026 and results describe those API versions without deployment safeguards
Methodology Notes
PsyArXiv preprint, DOI 10.31234/osf.io/3ep4v_v1, published 2026-05-15 (OSF date_published; version 1, modified 2026-08-13 with no version change), not peer-reviewed. OSF conflict-of-interest statement verbatim: 'Empathic Rocks, Inc. dba mpathic, is the sponsor of this research and the responsible party for its design, execution, and conclusions.' mpathic sells clinical-conversation AI products, so this is a vendor-sponsored benchmark. Affiliations from the PDF title block: mpathic plus UCSB Counseling, Clinical and School Psychology, UCSF Psychiatry and UW Psychiatry. Evaluator reliability: twelve candidates, nine cleared alpha 0.80 in calibration, one retained at 0.76 after review, ten evaluated. English only; six models; OSF records no data links. A companion eating-disorders benchmark (mPACT-ED-v1.0, DOI 10.31234/osf.io/htzcn_v1, same date and sponsor) reports a narrower 7.12 to 6.22 spread across the same six models. The 2026-09-03 watchlist entry that recorded only a DocSend-gated primary is resolved by these public preprints.
Topics
Authors
Alison Cerezo, Fuji Robledo Yamamoto, Alisa Breetz, Jasmin Acevedo, Shamal Lalvani, Caraline Bruzinski, Megan Greenlaw, Xieyining Huang, Danielle A. Schlosser, Sarah Peregrine Lord
Tags
Cite This
APA
Alison Cerezo et al. (2026). mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0). PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington). https://doi.org/10.31234/osf.io/3ep4v_v1
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School · 1 May 2026
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
arXiv (University of Aberdeen; University of Colorado Anschutz; Heriot-Watt University; University College London) · 1 Jun 2026
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
arXiv (MindSurf) · 29 Jun 2026