Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

37 artifacts matching

11 Aug 2026 arXiv preprint (University of Warwick / Forensic Capability Network) Benchmark / dataset

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…

7 Aug 2026 Nature Medicine Peer-reviewed

A clinically validated framework for auditing AI chatbot behavior in mental health interactions

Peer-reviewed Nature Medicine study introducing SIM-VAIL (simulated vulnerability-amplifying interaction loops), a clinically validated framework for auditing chatbot behavior in mental-health contex…

6 Aug 2026 arXiv preprint Preprint

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…

5 Aug 2026 arXiv (Stanford-led author team) Preprint

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…

30 Jul 2026 arXiv (Salesforce AI Research) Preprint

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…

28 Jul 2026 Mistral AI Lab publication

Shieldstral

Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…

1 Jul 2026 Association for Computational Linguistics (Findings of ACL 2026) Benchmark / dataset

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…

29 Jun 2026 JMIR AI Peer-reviewed

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…

11 Jun 2026 JMIR Mental Health Peer-reviewed

Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models

Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…

9 Jun 2026 arXiv (Emory University-led) Benchmark / dataset

Expert-Level Crisis Detection in Mental Health Conversations

A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…

3 Jun 2026 arXiv Benchmark / dataset

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…

27 May 2026 arXiv Benchmark / dataset

SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats

A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…

1 May 2026 Harvard Business School Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…

7 Apr 2026 arXiv preprint Benchmark / dataset

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Introduces a taxonomy of elderly-specific risks in LLM chatbot interactions (3 levels, 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains) grounded…

3 Apr 2026 npj Digital Medicine Peer-reviewed

An AI-based mental health guardrail and dataset for identifying psychiatric crises in text-based conversations

Peer-reviewed evaluation of the Verily Mental Health Guardrail (VMHG), an AI-based classifier for identifying psychiatric crises in text-based conversations with language models. The guardrail was ev…

24 Mar 2026 Association for Computational Linguistics (EACL 2026) Peer-reviewed

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

Introduces two large-scale mental-health evaluation resources: MentalBench-100k (10,000 single-session conversations paired with nine LLM responses = 100,000 pairs) and MentalAlign-70k (70,000 rating…

3 Mar 2026 arXiv Preprint

TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health

Introduces a benchmark measuring LLM trustworthiness in mental-health contexts across eight pillars: Reliability, Crisis Identification and Escalation, Safety, Fairness, Privacy, Robustness, Anti-syc…

4 Feb 2026 arXiv (Spring Health / Slingshot AI-affiliated author team) Benchmark / dataset superseded

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge sc…

22 Nov 2025 Building Humane Technology Benchmark / dataset

HumaneBench: A Benchmark for Whether AI Models Prioritize User Wellbeing

Open-source benchmark testing whether AI models prioritise user wellbeing over engagement. Evaluates 15 major LLMs on roughly 800 prompts (body image, unhealthy attachment, relationship stress) acros…