Skip to main content
Benchmark / dataset Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics

The first principle-grounded benchmark of LLM safety alignment in mental health built on a specific jurisdiction's professional ethics, the Australian psychology and psychiatry guidelines. It pairs multiple-choice ethical-knowledge items with open-ended behavioural tasks carrying fine-grained ethicality annotations. Across 14 models, refusal rates prove poor indicators of ethical behaviour, showing a divergence between safety triggers and clinical appropriateness, and several mental-health fine-tuned models underperform their base backbones in ethical alignment.

Publisher

Association for Computational Linguistics (Findings of ACL 2026); Monash University; University of Liverpool; Hefei University of Technology

Published

1 Jul 2026

Added

today

Key Findings

  • Refusal rates are poor indicators of ethical behaviour in mental-health contexts: clinically inadequate refusals can read as unempathetic and discourage help-seeking, and the benchmark shows safety triggers and clinical appropriateness diverging.
  • Domain-specific fine-tuning can degrade ethical robustness; several specialised mental-health models underperformed their base backbones in ethical alignment.
  • The benchmark grounds items in Australian guidelines (RANZCP and the Psychology Board of Australia per the beat's PDF read) and evaluates 14 models spanning mental-health LLMs, their base backbones and frontier comparators.
  • Open-ended responses carry fine-grained ethicality annotations (2,612 evaluated per the beat's PDF read) rather than a binary safe or unsafe label.

Methodology Notes

Benchmark construction from professional ethics guidelines, multiple-choice and open-ended tasks, 14 models evaluated with annotated ethicality; no patients. Published in Findings of ACL 2026 (July 2026), pages 39571 to 39589, DOI 10.18653/v1/2026.findings-acl.1971. Verified at the ACL Anthology on 2026-09-15; counts beyond the abstract come from the benchmarks beat's PDF read.

Authors

Yaling Shen, Stephanie Fong, Yiwen Jiang, Zimu Wang, Feilong Tang, Qingyang Xu, Xiangyu Zhao, Zhongxing Xu, Jiahe Liu, Jinpeng Hu, Dominic Dwyer, Zongyuan Ge

Tags

acl-2026mental-health-ethicsaustraliarefusalfine-tuning-degradationmonash

Cite This

APA

Yaling Shen et al. (2026). PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics. Association for Computational Linguistics (Findings of ACL 2026); Monash University; University of Liverpool; Hefei University of Technology. https://aclanthology.org/2026.findings-acl.1971/