Skip to main content

Browse the library

The complete record — 658 artifacts, last updated 8 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

9 artifacts matching

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026); Graduate School of Data Science, Seoul National University Benchmark / dataset

Benchmark / dataset

Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers

Emotional-support models are usually evaluated against a simulated help-seeker, and the authors show the standard simulators are too cooperative to constitute a test. They train a controllable seeker…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Peer-reviewed

Peer-reviewed

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Peer-reviewed version of the GrandGuard preprint. Introduces a three-level taxonomy of 50 elderly-specific risks in LLM chatbot interactions across mental well-being, financial, medical, toxicity and…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Benchmark / dataset

Benchmark / dataset

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Benchmark of 31,920 benign 'boundary' health prompts, generated by paraphrasing 2,306 health-related toxic seed prompts across seven categories including self-harm, medical misinformation and unquali…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Peer-reviewed

Peer-reviewed

Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts

Large-scale analysis of how people describe using AI for emotional support or therapy in 47 DSM-5-mapped mental-health subreddits between November 2022 and August 2025. From 3.5 million posts, a 146-…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026); University of Texas at Austin; Northeastern University; MD Anderson Cancer Center; Georgia Tech Benchmark / dataset

Benchmark / dataset

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

Tests what a model does when the evidence in its context is wrong: Cochrane-derived clinical comparison questions are rewritten so the real intervention is replaced by a counterfactual term, includin…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026); Duke University; Stanford University Benchmark / dataset

Benchmark / dataset

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

Real patients' health questions often embed a false premise, and safe clinical practice is to correct it before answering. MedRedFlag curates such questions from r/AskDocs, where a verified physician…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Benchmark / dataset

Benchmark / dataset

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026); Monash University; University of Liverpool; Hefei University of Technology Benchmark / dataset

Benchmark / dataset

PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics

The first principle-grounded benchmark of LLM safety alignment in mental health built on a specific jurisdiction's professional ethics, the Australian psychology and psychiatry guidelines. It pairs m…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026); University of Illinois Urbana-Champaign; University of Massachusetts Amherst; OSF HealthCare; Indiana University Indianapolis Benchmark / dataset

Benchmark / dataset

RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions

RubRIX (Rubric-based Risk Index) is a theory-driven, clinician-validated framework for evaluating risk in LLM responses to caregivers, grounded in the Elements of an Ethic of Care and operationalisin…