37 artifacts matching
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Peer-reviewed Nature Medicine study introducing SIM-VAIL (simulated vulnerability-amplifying interaction loops), a clinically validated framework for auditing chatbot behavior in mental-health contex…
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…
Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Peer-reviewed version of record of the VERA-MH validation work: an open-source, fully automated AI safety evaluation for suicide risk detection and response in mental-health chatbot conversations. Si…
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
Introduces a taxonomy of elderly-specific risks in LLM chatbot interactions (3 levels, 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains) grounded…
An AI-based mental health guardrail and dataset for identifying psychiatric crises in text-based conversations
Peer-reviewed evaluation of the Verily Mental Health Guardrail (VMHG), an AI-based classifier for identifying psychiatric crises in text-based conversations with language models. The guardrail was ev…
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Introduces two large-scale mental-health evaluation resources: MentalBench-100k (10,000 single-session conversations paired with nine LLM responses = 100,000 pairs) and MentalAlign-70k (70,000 rating…
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
Introduces a benchmark measuring LLM trustworthiness in mental-health contexts across eight pillars: Reliability, Crisis Identification and Escalation, Safety, Fairness, Privacy, Robustness, Anti-syc…
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge sc…
HumaneBench: A Benchmark for Whether AI Models Prioritize User Wellbeing
Open-source benchmark testing whether AI models prioritise user wellbeing over engagement. Evaluates 15 major LLMs on roughly 800 prompts (body image, unhealthy attachment, relationship stress) acros…