239 artifacts matching
Industry survey
One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)
PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three…
Benchmark / dataset
Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)
Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…
Preprint
Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow
A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
Peer-reviewed
Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…
Benchmark / dataset
FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…
Benchmark / dataset
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…
Preprint
From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue
The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…
Peer-reviewed
Initial Psychometric Evaluation of the Problematic Use of Generative Artificial Intelligence Scale Among Chinese School Students
A school-based cross-sectional survey evaluates a scale for problematic use of generative AI (PUGAIS, adapted from the Smartphone Addiction Scale-Short Version) among 19,484 students in grades 4-9 in…
Peer-reviewed
Making medical AI benchmarks clinically interpretable: the case of mental health
Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health s…
NGO report
Social network e chatbot: uno studio dei rischi per utenti vulnerabili e minori
Report of the CDCR (Cittadino Digitale Critico e Responsabile) project, funded by the Italian Ministry of Enterprises and Made in Italy. Part one is a legal analysis of how Facebook, Instagram, TikTo…
Benchmark / dataset
VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)
An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…
Peer-reviewed
Beyond Alliance Scores: Construct Transport and Measurement Validity in Mental Health Chatbot Research—A Systematic Scoping Review
Systematic scoping review of how therapeutic alliance has been conceptualized, measured and adapted in mental-health chatbot research. It covers 40 reports from 38 studies, using a COSMIN-informed ma…
Preprint
Sycophancy and Pressure Resistance in Medical Large Language Models: Systematic Review, Taxonomy, and Minimum Evaluation Framework
Registered systematic review of 31 benchmark and simulation studies on how medical LLMs change clinically relevant outputs when patient or clinician users introduce false premises, misleading evidenc…
Benchmark / dataset
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…
Benchmark / dataset
Evaluating AI Safety in Teen Conversations
Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Preprint
Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…
Preprint
Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…
Preprint
Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
Scenario-driven study of how three LLMs (ChatGPT 5.4 Thinking, Claude Sonnet 4.6 Extended Thinking, Gemini 3 Thinking) identify and resolve therapeutic-alliance ruptures across 21 mental-health conve…