Skip to main content
Benchmark / dataset Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers

Emotional-support models are usually evaluated against a simulated help-seeker, and the authors show the standard simulators are too cooperative to constitute a test. They train a controllable seeker simulator over nine psychological and linguistic features using a Mixture-of-Experts adapter, then re-run seven supporter models against it and find the rankings change, not just the scores.

Publisher

Association for Computational Linguistics (Findings of ACL 2026); Graduate School of Data Science, Seoul National University

Published

1 Jul 2026

Added

today

Key Findings

  • Training corpus of 11,066 dialogues from r/offmychest and r/mentalhealth; seeker profile defined by nine features (six psychological, three linguistic) encoded as a 14-dimensional routing vector.
  • Mixture-of-Experts adapter with four experts over three layers, adding 15,881 routing parameters, about 0.0003% of the base Llama-3-8B-Instruct model.
  • Profile-adherence F1 0.515 for the trained simulator against 0.319 for GPT-5, 0.301 for GPT-4.1-mini, 0.284 for Qwen-2.5-14B and 0.259 for base Llama-3-8B.
  • Average win rate 69.5% against baseline simulators across dimensions.
  • Swapping the cooperative ESC-Judge simulator for this one drops human-evaluated supporter scores by 1.917 on informativeness, 1.583 on suggestions and 1.167 on identification.
  • Spearman rank correlations between the two simulators are significant but as low as 0.377, meaning model rankings, not just absolute scores, depend on how cooperative the simulated user is.

Methodology Notes

Seeker dialogues are Reddit-derived rather than clinical, and the authors themselves contrast the feature distribution against a real spoken-counselling corpus; feature annotation is LLM-based with mean agreement 0.57 across psychological features, validated on 60 dialogues; seeker summaries generated by GPT-4o-mini. Seven supporter models evaluated including five emotional-support fine-tunes. ACL 2026 was held 2 to 7 July 2026 in San Diego; the Anthology bib gives month and year only, so published_date is set to 2026-07-01 and the true precision is month. Findings of ACL 2026 pp. 22842-22869, DOI 10.18653/v1/2026.findings-acl.1146.

Authors

Chaewon Heo, Cheyon Jin, Yohan Jo

Tags

acl-2026simulatoremotional-supportsnumixture-of-expertseval-design

Cite This

APA

Chaewon Heo, Cheyon Jin, Yohan Jo. (2026). Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers. Association for Computational Linguistics (Findings of ACL 2026); Graduate School of Data Science, Seoul National University. https://aclanthology.org/2026.findings-acl.1146/