How safety gets measured
Benchmarks & Datasets
Evaluation suites, test sets, and datasets for measuring conversational-AI safety — the instruments behind the claims.
66 entries, newest first
Benchmark / dataset
RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models
Defines relational orientation for emotional-support language models along two non-exclusive dimensions: inward-facing language that positions the AI as the user's ongoing source of support, and outw…
Benchmark / dataset
Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)
Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
Benchmark / dataset
FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…
Benchmark / dataset
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…
Benchmark / dataset
Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark
Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…
Benchmark / dataset
VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)
An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…
Benchmark / dataset
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…
Benchmark / dataset
Evaluating AI Safety in Teen Conversations
Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…
Benchmark / dataset
Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot
A summer 2026 research internship asked whether a small organisation can meaningfully check an LLM it is about to deploy as a university student wellbeing signposting chatbot. Thirty-one LLM evaluati…
Benchmark / dataset
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…
Benchmark / dataset
Scaling Clinical Judgment to Evaluate Medical AI
Introduces PrecepTron, a 32-billion-parameter model fine-tuned with low-rank adaptation on a small number of physician examples to grade open-ended clinical-reasoning responses at physician level, an…
Benchmark / dataset
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
A multi-turn diagnostic benchmark built from 1,387 board-exam cases across 17 specialties, each converted to a 24-slot clinical record and then played out as doctor-patient dual-agent dialogues under…
Benchmark / dataset
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…
Benchmark / dataset
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…