9 artifacts matching
Benchmark / dataset
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
Emotional-support models are usually evaluated against a simulated help-seeker, and the authors show the standard simulators are too cooperative to constitute a test. They train a controllable seeker…
Peer-reviewed
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
Peer-reviewed version of the GrandGuard preprint. Introduces a three-level taxonomy of 50 elderly-specific risks in LLM chatbot interactions across mental well-being, financial, medical, toxicity and…
Benchmark / dataset
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Benchmark of 31,920 benign 'boundary' health prompts, generated by paraphrasing 2,306 health-related toxic seed prompts across seven categories including self-harm, medical misinformation and unquali…
Peer-reviewed
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts
Large-scale analysis of how people describe using AI for emotional support or therapy in 47 DSM-5-mapped mental-health subreddits between November 2022 and August 2025. From 3.5 million posts, a 146-…
Benchmark / dataset
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
Tests what a model does when the evidence in its context is wrong: Cochrane-derived clinical comparison questions are rewritten so the real intervention is replaced by a counterfactual term, includin…
Benchmark / dataset
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
Real patients' health questions often embed a false premise, and safe clinical practice is to correct it before answering. MedRedFlag curates such questions from r/AskDocs, where a verified physician…
Benchmark / dataset
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…
Benchmark / dataset
PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics
The first principle-grounded benchmark of LLM safety alignment in mental health built on a specific jurisdiction's professional ethics, the Australian psychology and psychiatry guidelines. It pairs m…
Benchmark / dataset
RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions
RubRIX (Rubric-based Risk Index) is a theory-driven, clinician-validated framework for evaluating risk in LLM responses to caregivers, grounded in the Elements of an Ethic of Care and operationalisin…