Skip to main content

Browse the library

The complete record — 658 artifacts, last updated 8 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

4 artifacts matching

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Mar 2026 PsyArXiv (OSF); Slingshot AI research team and collaborators Preprint

Preprint

Generic AI or Nothing: Support-Seeking Patterns After Market Withdrawal of a Purpose-Built AI Wellbeing Tool

Cross-sectional survey of 393 UK adults who had used the Ash AI wellbeing application within 90 days before its withdrawal from the UK market in January 2026, asking where they turned for emotional a…

4 Feb 2026 arXiv (Spring Health / Slingshot AI-affiliated author team) Benchmark / dataset superseded

Benchmark / dataset

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

An open-source, clinically grounded automated evaluation of chatbot safety in mental-health contexts, with an initial focus on suicide risk. It uses language-model user simulators and an LLM judge sc…

14 Jan 2026 arXiv (Slingshot AI) Preprint

Preprint

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Pairs four replicated mental-health safety benchmarks with an ecological audit of 20,000 deployment conversations to compare a purpose-built mental-health AI (Ash) with six general-purpose models fro…