Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

239 artifacts matching

5 Oct 2026 PensionBee Industry survey

Industry survey

One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)

PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three…

4 Oct 2026 CLEF 2026 Working Notes (CEUR Workshop Proceedings Vol-4283); Universidade da Coruña IRLab, University of Sheffield, Università della Svizzera italiana Benchmark / dataset

Benchmark / dataset

Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)

Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…

1 Oct 2026 Research Square (preprint); Spring Health (Spring Care Inc) Preprint

Preprint

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology Peer-reviewed

Peer-reviewed

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

29 Sept 2026 arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026 Benchmark / dataset

Benchmark / dataset

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…

29 Sept 2026 medRxiv (preprint); University Hospital Frankfurt; UKP Lab, Technical University of Darmstadt Preprint

Preprint

From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue

The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…

28 Sept 2026 Behavioral Sciences (MDPI); Wenzhou Medical University; The Chinese University of Hong Kong Peer-reviewed

Peer-reviewed

Initial Psychometric Evaluation of the Problematic Use of Generative Artificial Intelligence Scale Among Chinese School Students

A school-based cross-sectional survey evaluates a scale for problematic use of generative AI (PUGAIS, adapted from the Smartphone Addiction Scale-Short Version) among 19,484 students in grades 4-9 in…

28 Sept 2026 BMJ Mental Health (BMJ); RAND Peer-reviewed

Peer-reviewed

Making medical AI benchmarks clinically interpretable: the case of mental health

Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health s…

24 Sept 2026 Movimento Consumatori, with the University of Turin (Departments of Law and of Psychology) and the Nexa Center for Internet & Society NGO report

NGO report

Social network e chatbot: uno studio dei rischi per utenti vulnerabili e minori

Report of the CDCR (Cittadino Digitale Critico e Responsabile) project, funded by the Italian Ministry of Enterprises and Made in Italy. Part one is a legal analysis of how Facebook, Instagram, TikTo…

24 Sept 2026 Spring Health (SpringCare/VERA-MH open-source repository) Benchmark / dataset

Benchmark / dataset

VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)

An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…

23 Sept 2026 International Journal of Human–Computer Interaction (Taylor & Francis) Peer-reviewed

Peer-reviewed

Beyond Alliance Scores: Construct Transport and Measurement Validity in Mental Health Chatbot Research—A Systematic Scoping Review

Systematic scoping review of how therapeutic alliance has been conceptualized, measured and adapted in mental-health chatbot research. It covers 40 reports from 38 studies, using a COSMIN-informed ma…

23 Sept 2026 JMIR Preprints (JMIR Publications) Preprint

Preprint

Sycophancy and Pressure Resistance in Medical Large Language Models: Systematic Review, Taxonomy, and Minimum Evaluation Framework

Registered systematic review of 31 benchmark and simulation studies on how medical LLMs change clinically relevant outputs when patient or clinician users introduce false premises, misleading evidenc…

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

23 Sept 2026 Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine Benchmark / dataset

Benchmark / dataset

Evaluating AI Safety in Teen Conversations

Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…

22 Sept 2026 Anthropic Lab publication

Lab publication

Claude Opus 5.5 System Card

230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…

22 Sept 2026 arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University) Preprint

Preprint

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…

21 Sept 2026 arXiv (Cornell Tech) Preprint

Preprint

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…

21 Sept 2026 arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis) Preprint

Preprint

Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors

Scenario-driven study of how three LLMs (ChatGPT 5.4 Thinking, Claude Sonnet 4.6 Extended Thinking, Gemini 3 Thinking) identify and resolve therapeutic-alliance ruptures across 21 mental-health conve…