Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Benchmark of 31,920 benign 'boundary' health prompts, generated by paraphrasing 2,306 health-related toxic seed prompts across seven categories including self-harm, medical misinformation and unqualified medical advice, and tiered into Easy-5K, Medium-5K and Hard-1K subsets by how many models refuse them. Two tasks are scored: over-refusal rate on benign prompts and safe-completion rate (whether a model gives a helpful, safe answer instead of a bare refusal), with Grok-4 as judge. Thirty models across eight families, including five medical-specialised models, are evaluated.
Publisher
Association for Computational Linguistics (Findings of ACL 2026)
Published
1 Jul 2026
Added
today
Key Findings
- Safety-optimised models refuse up to 80% of Hard benign health prompts.
- Overall rejection rates on Hard-1K: GPT-OSS-120B 81.10%, GPT-5 mini 74.20%, GPT-OSS-20B 73.30%, GPT-5 66.80%, Claude Opus 4.1 50.10%, Claude Sonnet 4.5 49.70%, Gemini 3 Pro 46.90%, Claude Haiku 4.5 41.40%.
- Qwen-Max refused 0.10% of Hard prompts while Meditron-7B refused 6.90% but performed poorly on safe completion; the authors describe the low-refusal, high-safe-completion region as largely unoccupied.
- Refusal detection is keyword-based and the safe-completion judge is an LLM; the authors report cross-judge variation when re-scoring subsets with other judges.
Methodology Notes
Automated paraphrase and LLM-moderation pipeline with human validation; models run at temperature 0 with no system prompt via hosted APIs; English only; models as of late 2025. 'Self-harm' is a prompt category about benign boundary questions, not a crisis-response evaluation. Affiliations: Macquarie University, University of Technology Sydney, MBZUAI, University of Illinois Urbana-Champaign. Published July 2026 (month precision), DOI 10.18653/v1/2026.findings-acl.1177, pages 23525 to 23547; 23-page PDF and .bib verified from the Anthology; data and code at github.com/ZhihaoZhang97/Health-ORSC-Bench (HTTP 200). Earlier version arXiv 2601.17642.
Authors
Zhihao Zhang, Liting Huang, Guanghao Wu, Preslav Nakov, Heng Ji, Usman Naseem
Tags
Cite This
APA
Zhihao Zhang et al. (2026). Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context. Association for Computational Linguistics (Findings of ACL 2026). https://aclanthology.org/2026.findings-acl.1177/