Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

120 artifacts matching

4 Oct 2026 CLEF 2026 Working Notes (CEUR Workshop Proceedings Vol-4283); Universidade da Coruña IRLab, University of Sheffield, Università della Svizzera italiana Benchmark / dataset

Benchmark / dataset

Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)

Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

29 Sept 2026 arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026 Benchmark / dataset

Benchmark / dataset

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…

28 Sept 2026 arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University) Benchmark / dataset

Benchmark / dataset

Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark

Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…

28 Sept 2026 BMJ Mental Health (BMJ); RAND Peer-reviewed

Peer-reviewed

Making medical AI benchmarks clinically interpretable: the case of mental health

Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health s…

24 Sept 2026 Spring Health (SpringCare/VERA-MH open-source repository) Benchmark / dataset

Benchmark / dataset

VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)

An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

23 Sept 2026 Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine Benchmark / dataset

Benchmark / dataset

Evaluating AI Safety in Teen Conversations

Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…

21 Sept 2026 arXiv (Cornell Tech) Preprint

Preprint

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…

18 Sept 2026 University of Nottingham, School of Computer Science (Responsible AI UK Cornerstone 2 AI Assurance programme) Benchmark / dataset

Benchmark / dataset

Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot

A summer 2026 research internship asked whether a small organisation can meaningfully check an LLM it is about to deploy as a university student wellbeing signposting chatbot. Thirty-one LLM evaluati…

17 Sept 2026 arXiv (Georgia Institute of Technology) Preprint

Preprint

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

Introduces COPES (Community-centered Peer Engaged Support), a dataset of mental-health support-seeking Reddit queries with community-endorsed responses (5,536 posts after filtering, across five commu…

14 Sept 2026 arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut) Benchmark / dataset

Benchmark / dataset

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…

11 Sept 2026 arXiv (Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford; Massachusetts General Hospital; University of Alberta; MIT; Erasmus MC; University of Maryland) Benchmark / dataset

Benchmark / dataset

Scaling Clinical Judgment to Evaluate Medical AI

Introduces PrecepTron, a 32-billion-parameter model fine-tuned with low-rank adaptation on a small number of physician examples to grade open-ended clinical-reasoning responses at physician level, an…

11 Sept 2026 arXiv (MBZUAI; Cairo University; Ain Shams University; CSIRO; INSAIT, Sofia University) Benchmark / dataset

Benchmark / dataset

MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

A multi-turn diagnostic benchmark built from 1,387 board-exam cases across 17 specialties, each converted to a 24-slot clinical record and then played out as doctor-patient dual-agent dialogues under…

11 Sept 2026 arXiv (University of Oxford) Preprint

Preprint

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

Tests whether rubric-based evaluation, the dominant approach for grading LLMs in medicine, detects clinically relevant hallucinations. After a controlled study on MedHallu showing more specific rubri…

11 Sept 2026 Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology Peer-reviewed

Peer-reviewed

Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage

A blinded, paired benchmark of three consumer assistants answering patient- and caregiver-facing questions about late-life depression. Ninety questions covering six geriatric-psychiatry domains, stra…

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 arXiv (Stanford University) Preprint

Preprint

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…

8 Sept 2026 arXiv (University of Michigan; Stanford University; Yale University; Microsoft Research; Abridge; Cornell Tech) Preprint

Preprint

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

The paper adapts convergent and discriminant validity from the social sciences into a procedure for interrogating whether AI benchmarks measure the concepts they claim to measure, applying it to 56 c…