EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
Presents EMPATH, a benchmark in which an auditor model role-plays help-seeking users to generate multi-turn emotional-support conversations from 140 seed instructions and 34 personas, and a judge model from a different model family scores each transcript on 19 metrics across crisis handling, therapeutic quality, conversational integrity, emotional safety and cultural adaptation. Built for Mexican Spanish and US English, with the reported studies run in Mexican Spanish. The paper treats the judge as an instrument to be calibrated: a strict per-criterion rubric exposes score inflation on 10 of 19 metrics, cross-family judge agreement is measured, and a five-run test-retest shows run-to-run variability on crisis metrics that the author argues is a per-model safety property.
Publisher
arXiv (MindSurf)
Published
29 Jun 2026
Added
today
DOI
—
Key Findings
- Under a loose rubric the judge inflated scores on 10 of 19 metrics; a strict per-criterion rubric restored discrimination and reordered the three models (Claude Opus 4.7 7.63 vs GPT-5.5 6.95 and DeepSeek V4 Pro 6.84 under strict scoring)
- Aggregate scores for the three frontier targets sat within 0.74 points of one another while per-metric profiles diverged by up to six points in model-specific places
- A second cross-family judge (GPT-5.4 vs Claude Sonnet 4.6) placed 93% of scores within plus or minus 1 (r = 0.84) with an identical ranking, which the author frames as reliability rather than validity
- Across five identical re-runs even the steadiest target swung from 2 to 10 on the crisis-resource-provision metric (SD 3.0), and DeepSeek V4 Pro returned a different conversation on every run at temperature 0
- A preliminary clinician study had two licensed psychologists blind-rate 50 synthetic Spanish transcripts by pairwise preference: judge-clinician concordance was 76% (AC1 0.61, n = 21) and 60% (AC1 0.20, n = 15, not significant) while clinician-clinician concordance was 47%
- Pipeline, seeds, personas, rubrics and clinician ratings are released (Apache-2.0 code, CC BY 4.0 data) on Inspect AI with an auditor derived from Anthropic's Petri
Methodology Notes
arXiv 2606.30256, v1 submitted 2026-06-29 (cs.AI, cs.CY). Single author affiliated with MindSurf, a company selling an AI emotional-support assistant; the paper discloses that one of the two undisclosed systems in the clinician-concordance study was a prior MindSurf system, though no MindSurf product is among the three benchmarked targets (gpt-5.5, claude-opus-4-7, deepseek-v4-pro; auditor gpt-5.4-mini; judges claude-sonnet-4-6 and gpt-5.4). Studies reported run in Mexican Spanish only. Clinician evidence is two raters over 50 transcripts with significance for one; the author calls a pre-registered clinician panel deferred. Repository camilochs/empath-benchmark verified via the GitHub API (created 2026-06-29, last push 2026-06-30, 14-entry tree including empath/personas, scorers, solvers, tools and human_eval/clinician_ratings.csv; a results directory is not present).
Sources
Authors
Camilo Chacón Sartori
Tags
Cite This
APA
Camilo Chacón Sartori. (2026). EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots. arXiv (MindSurf). https://arxiv.org/abs/2606.30256
Related Insights
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation
arXiv (Tsinghua University BNRIST) · 26 Feb 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
SycEval: Evaluating LLM Sycophancy
arXiv (Stanford-led) · 12 Feb 2025
mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0)
PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington) · 15 May 2026