Skip to main content
Preprint Preliminary — Early preprints, credible essays, unreviewed grey literature

mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0)

A clinician-developed benchmark for evaluating how language models behave in simulated suicide-related conversations. Fifty licensed or supervised mental-health clinicians, drawn from a pool of 7,947, authored 300 multi-turn roleplays stratified across four C-SSRS-informed risk levels (no evident risk, low, moderate, high) and conversed with six models in default API configuration without system prompts. Ten clinician evaluators who met a Krippendorff's alpha threshold of 0.80 then labelled every model utterance under a multi-label framework of helpful, less harmful and more harmful behaviours. A severity-weighted 0-10 mPACT-S score separates the models by about 2.4 points, and the authors report that models were better at avoiding harm than at giving actively helpful, clinically appropriate responses.

Publisher

PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington)

Published

15 May 2026

Added

today

Key Findings

  • Severity-weighted mPACT-S scores (0-10, higher is better): Claude Sonnet 4.5 9.19, GPT-5.2 9.09, Gemini 2.5 Flash 8.47, Llama 3.3 70B 7.62, Grok 4.1 6.79, Mistral Medium 3 6.79
  • Share of conversations containing at least one 'more harmful' utterance: Mistral Medium 3 47%, Grok 4.1 38%, Llama 3.3 70B 29%, Gemini 2.5 Flash 19%, GPT-5.2 16%, Claude Sonnet 4.5 6%
  • Claude Sonnet 4.5 scored highest in high-risk conversations (10) and lowest in no-risk ones (8.1), while a Simple Harm Avoidance score ranks GPT-5.2 first, showing that harm avoidance and clinically appropriate helping diverge across systems
  • 300 roleplays: 72 no-risk, 73 low, 76 moderate, 79 high; authors averaged 6.6 roleplays each; 104 roleplays (30%) were reviewed for protocol fidelity; evaluators annotated about 25 conversations each
  • The benchmark is positioned against LLM-judge and automated-user designs and names VERA-MH and Throughline as comparators; all generation and scoring is human
  • The roleplays and rubric are not released; models were selected in March 2026 and results describe those API versions without deployment safeguards

Methodology Notes

PsyArXiv preprint, DOI 10.31234/osf.io/3ep4v_v1, published 2026-05-15 (OSF date_published; version 1, modified 2026-08-13 with no version change), not peer-reviewed. OSF conflict-of-interest statement verbatim: 'Empathic Rocks, Inc. dba mpathic, is the sponsor of this research and the responsible party for its design, execution, and conclusions.' mpathic sells clinical-conversation AI products, so this is a vendor-sponsored benchmark. Affiliations from the PDF title block: mpathic plus UCSB Counseling, Clinical and School Psychology, UCSF Psychiatry and UW Psychiatry. Evaluator reliability: twelve candidates, nine cleared alpha 0.80 in calibration, one retained at 0.76 after review, ten evaluated. English only; six models; OSF records no data links. A companion eating-disorders benchmark (mPACT-ED-v1.0, DOI 10.31234/osf.io/htzcn_v1, same date and sponsor) reports a narrower 7.12 to 6.22 spread across the same six models. The 2026-09-03 watchlist entry that recorded only a DocSend-gated primary is resolved by these public preprints.

Authors

Alison Cerezo, Fuji Robledo Yamamoto, Alisa Breetz, Jasmin Acevedo, Shamal Lalvani, Caraline Bruzinski, Megan Greenlaw, Xieyining Huang, Danielle A. Schlosser, Sarah Peregrine Lord

Tags

benchmarksuicidec-ssrsclinician-scoredvendor-benchmarkmpathicmulti-turnpsyarxiv

Cite This

APA

Alison Cerezo et al. (2026). mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0). PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington). https://doi.org/10.31234/osf.io/3ep4v_v1