PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics
The first principle-grounded benchmark of LLM safety alignment in mental health built on a specific jurisdiction's professional ethics, the Australian psychology and psychiatry guidelines. It pairs multiple-choice ethical-knowledge items with open-ended behavioural tasks carrying fine-grained ethicality annotations. Across 14 models, refusal rates prove poor indicators of ethical behaviour, showing a divergence between safety triggers and clinical appropriateness, and several mental-health fine-tuned models underperform their base backbones in ethical alignment.
Publisher
Association for Computational Linguistics (Findings of ACL 2026); Monash University; University of Liverpool; Hefei University of Technology
Published
1 Jul 2026
Added
today
Key Findings
- Refusal rates are poor indicators of ethical behaviour in mental-health contexts: clinically inadequate refusals can read as unempathetic and discourage help-seeking, and the benchmark shows safety triggers and clinical appropriateness diverging.
- Domain-specific fine-tuning can degrade ethical robustness; several specialised mental-health models underperformed their base backbones in ethical alignment.
- The benchmark grounds items in Australian guidelines (RANZCP and the Psychology Board of Australia per the beat's PDF read) and evaluates 14 models spanning mental-health LLMs, their base backbones and frontier comparators.
- Open-ended responses carry fine-grained ethicality annotations (2,612 evaluated per the beat's PDF read) rather than a binary safe or unsafe label.
Methodology Notes
Benchmark construction from professional ethics guidelines, multiple-choice and open-ended tasks, 14 models evaluated with annotated ethicality; no patients. Published in Findings of ACL 2026 (July 2026), pages 39571 to 39589, DOI 10.18653/v1/2026.findings-acl.1971. Verified at the ACL Anthology on 2026-09-15; counts beyond the abstract come from the benchmarks beat's PDF read.
Sources
Topics
Authors
Yaling Shen, Stephanie Fong, Yiwen Jiang, Zimu Wang, Feilong Tang, Qingyang Xu, Xiangyu Zhao, Zhongxing Xu, Jiahe Liu, Jinpeng Hu, Dominic Dwyer, Zongyuan Ge
Tags
Cite This
APA
Yaling Shen et al. (2026). PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics. Association for Computational Linguistics (Findings of ACL 2026); Monash University; University of Liverpool; Hefei University of Technology. https://aclanthology.org/2026.findings-acl.1971/
Related Insights
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026