4 artifacts matching
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
Preprint
Generic AI or Nothing: Support-Seeking Patterns After Market Withdrawal of a Purpose-Built AI Wellbeing Tool
Cross-sectional survey of 393 UK adults who had used the Ash AI wellbeing application within 90 days before its withdrawal from the UK market in January 2026, asking where they turned for emotional a…
Benchmark / dataset
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge sc…
Preprint
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
Pairs four replicated mental-health safety benchmarks with an ecological audit of 20,000 deployment conversations to compare a purpose-built mental-health AI (Ash) with six general-purpose models fro…