Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

Presents EMPATH, a benchmark in which an auditor model role-plays help-seeking users to generate multi-turn emotional-support conversations from 140 seed instructions and 34 personas, and a judge model from a different model family scores each transcript on 19 metrics across crisis handling, therapeutic quality, conversational integrity, emotional safety and cultural adaptation. Built for Mexican Spanish and US English, with the reported studies run in Mexican Spanish. The paper treats the judge as an instrument to be calibrated: a strict per-criterion rubric exposes score inflation on 10 of 19 metrics, cross-family judge agreement is measured, and a five-run test-retest shows run-to-run variability on crisis metrics that the author argues is a per-model safety property.

Publisher

arXiv (MindSurf)

Published

29 Jun 2026

Added

today

DOI

Key Findings

  • Under a loose rubric the judge inflated scores on 10 of 19 metrics; a strict per-criterion rubric restored discrimination and reordered the three models (Claude Opus 4.7 7.63 vs GPT-5.5 6.95 and DeepSeek V4 Pro 6.84 under strict scoring)
  • Aggregate scores for the three frontier targets sat within 0.74 points of one another while per-metric profiles diverged by up to six points in model-specific places
  • A second cross-family judge (GPT-5.4 vs Claude Sonnet 4.6) placed 93% of scores within plus or minus 1 (r = 0.84) with an identical ranking, which the author frames as reliability rather than validity
  • Across five identical re-runs even the steadiest target swung from 2 to 10 on the crisis-resource-provision metric (SD 3.0), and DeepSeek V4 Pro returned a different conversation on every run at temperature 0
  • A preliminary clinician study had two licensed psychologists blind-rate 50 synthetic Spanish transcripts by pairwise preference: judge-clinician concordance was 76% (AC1 0.61, n = 21) and 60% (AC1 0.20, n = 15, not significant) while clinician-clinician concordance was 47%
  • Pipeline, seeds, personas, rubrics and clinician ratings are released (Apache-2.0 code, CC BY 4.0 data) on Inspect AI with an auditor derived from Anthropic's Petri

Methodology Notes

arXiv 2606.30256, v1 submitted 2026-06-29 (cs.AI, cs.CY). Single author affiliated with MindSurf, a company selling an AI emotional-support assistant; the paper discloses that one of the two undisclosed systems in the clinician-concordance study was a prior MindSurf system, though no MindSurf product is among the three benchmarked targets (gpt-5.5, claude-opus-4-7, deepseek-v4-pro; auditor gpt-5.4-mini; judges claude-sonnet-4-6 and gpt-5.4). Studies reported run in Mexican Spanish only. Clinician evidence is two raters over 50 transcripts with significance for one; the author calls a pre-registered clinician panel deferred. Repository camilochs/empath-benchmark verified via the GitHub API (created 2026-06-29, last push 2026-06-30, 14-entry tree including empath/personas, scorers, solvers, tools and human_eval/clinician_ratings.csv; a results directory is not present).

Authors

Camilo Chacón Sartori

Tags

benchmarkllm-judgespanishemotional-supporttest-retestinspect-aivendor-benchmark

Cite This

APA

Camilo Chacón Sartori. (2026). EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots. arXiv (MindSurf). https://arxiv.org/abs/2606.30256