228 artifacts matching
Benchmark / dataset
Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)
Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…
Peer-reviewed
Artificial Intelligence–Associated Psychosis
A psychiatric case report from the mental health division of Northern Hospital, Melbourne, describes a 27-year-old woman with a year of psychotic symptoms whose presentation centred on ChatGPT. She u…
Preprint
Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow
A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
NGO report
Generative AI Use and Mind Health Outcomes
Sapien Labs rapid report analysing generative AI chatbot use among 264,085 adults in the Global Mind Project's online survey across 85+ countries, relating frequency and purpose of use to the Mind He…
Peer-reviewed
AI or Human Support for Suicide Prevention? Examining Help-Seeking Intention in Suicidal Crisis
A cross-sectional web survey of 1,024 Japanese adults aged 18-69 with severe psychological distress (Kessler-6 score of 13 or more) asks whether they would use chat-based crisis support delivered by…
Preprint
How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan
Data-donation study of complete ChatGPT exports from 1,252 users in India, Nigeria, Brazil and Pakistan (202,590 conversations, December 2022 to February 2026), paired with self-reported age and gend…
Peer-reviewed
AI-Mediated Mental Health Support: The Role of Attachment Orientation and Psychological Distress
A preregistered cross-sectional survey of 584 Israeli adults who use general-purpose generative AI asks whether attachment orientation (ECR-RS) and psychological distress (DASS-21) are associated wit…
Preprint
Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT
Mixed-methods study of how young adults aged 18 to 25 use ChatGPT when distressed, built on complete donated ChatGPT histories (19,930 conversations from 158 participants) plus a survey that included…
Peer-reviewed
Initial Psychometric Evaluation of the Problematic Use of Generative Artificial Intelligence Scale Among Chinese School Students
A school-based cross-sectional survey evaluates a scale for problematic use of generative AI (PUGAIS, adapted from the Smartphone Addiction Scale-Short Version) among 19,484 students in grades 4-9 in…
Peer-reviewed
Making medical AI benchmarks clinically interpretable: the case of mental health
Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health s…
Industry survey
NPR Religion and AI Survey (Ipsos): Most Americans skeptical of turning to AI for spiritual guidance
Ipsos KnowledgePanel probability survey of 1,011 US adults for NPR on the use of AI tools for personal, emotional and spiritual counsel. It reports how often AI users seek advice on personal decision…
Peer-reviewed
Beyond Replacement: Episode-Based Pathways Between Generative AI and School Counselor Help-Seeking Among Indonesian Adolescents
Explanatory sequential mixed-methods study of 432 Indonesian students aged 15 to 18. It maps how GenAI and human support, including school counselors, are used within the same personal or emotional p…
Benchmark / dataset
VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)
An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…
Peer-reviewed
One Question, Four Voices: How Advice for Alzheimer's Caregiving Differs Between Caregivers, Clinicians, and Large Language Models
A side-by-side comparison of responses to 85 real caregiver questions from the ALZConnected forum across peer caregivers, physicians, ChatGPT and CareGPT, a retrieval-augmented variant grounded in pe…
Peer-reviewed
Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions
Compares session-level behaviour of human peer counsellors and an LLM counsellor trained on the same single-session CBT manual. An 18-month ethnography on a peer-support platform (seven counsellors,…
Peer-reviewed
Parental Attachment Anxiety and Adolescents' Authentic Self-Disclosure to Generative AI: The Roles of Rumination, Depression, and Gender
Offline survey of 685 Chinese adolescents (324 girls; mean age 13.17) testing how parent-adolescent attachment anxiety relates to adolescents' authentic self-disclosure to generative AI, drawing on c…
Peer-reviewed
Beyond Alliance Scores: Construct Transport and Measurement Validity in Mental Health Chatbot Research—A Systematic Scoping Review
Systematic scoping review of how therapeutic alliance has been conceptualized, measured and adapted in mental-health chatbot research. It covers 40 reports from 38 studies, using a COSMIN-informed ma…
Benchmark / dataset
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…
Peer-reviewed
Language and cultural prompts influence large language models' responses to culturally anchored mental health attitudes
Two factorial experiments (2 prompt languages, Chinese or English, by 2 cultural framings, China or the United States) presenting ChatGPT-4o and DeepSeek-V3 with the same description of a friend show…