642 artifacts
Industry survey
One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)
PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three…
NGO report
Sicher mit KI aufwachsen. Zwischen Algorithmen, Chatbots und Deepfakes: Schutz, Teilhabe und Bildung für Kinder und Jugendliche (Vorstudie)
A 67-page German-language pre-study commissioned by the child-rights organisations SOS-Kinderdörfer weltweit and Terre des Hommes and written with Klartext AI. It maps where children encounter AI, th…
Benchmark / dataset
Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)
Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…
Preprint
Mitigating Social Sycophancy via Pluralistic Preference Optimization
Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…
Government report
AI and complaints: removing barriers, reinforcing divides? How AI is influencing who complains, how they complain, and what could come next
Research briefing from the Legal Ombudsman for England and Wales on how consumers use generative AI when deciding whether and how to complain about regulated services. It combines an Ipsos survey of…
Peer-reviewed
Artificial Intelligence–Associated Psychosis
A psychiatric case report from the mental health division of Northern Hospital, Melbourne, describes a 27-year-old woman with a year of psychotic symptoms whose presentation centred on ChatGPT. She u…
Preprint
Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow
A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
Peer-reviewed
Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…
Benchmark / dataset
FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…
Preprint
Persona and Persuasive Framing in AI Voice Agents: A 2x2 Field Experiment with Children
Randomised 2x2 field experiment embedded in a public German-language Santa Claus telephone hotline before Christmas 2025, in which children's calls were routed to LLM voice agents that varied persona…
Industry survey
Generatieve AI en illegaal online gokaanbod: Hoe AI-tools Nederlandse consumenten bij niet-vergunde casino's brengen
Dutch-language test of ten consumer generative AI tools on whether simple, realistic questions lead users to online casinos without a Dutch licence. Each tool received nine prompts in three series: n…
NGO report
Generative AI Use and Mind Health Outcomes
Sapien Labs rapid report analysing generative AI chatbot use among 264,085 adults in the Global Mind Project's online survey across 85+ countries, relating frequency and purpose of use to the Mind He…
Peer-reviewed
AI or Human Support for Suicide Prevention? Examining Help-Seeking Intention in Suicidal Crisis
A cross-sectional web survey of 1,024 Japanese adults aged 18-69 with severe psychological distress (Kessler-6 score of 13 or more) asks whether they would use chat-based crisis support delivered by…
Preprint
How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan
Data-donation study of complete ChatGPT exports from 1,252 users in India, Nigeria, Brazil and Pakistan (202,590 conversations, December 2022 to February 2026), paired with self-reported age and gend…
Benchmark / dataset
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…
Peer-reviewed
AI-Mediated Mental Health Support: The Role of Attachment Orientation and Psychological Distress
A preregistered cross-sectional survey of 584 Israeli adults who use general-purpose generative AI asks whether attachment orientation (ECR-RS) and psychological distress (DASS-21) are associated wit…
Preprint
From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue
The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…
Lab publication
System Card: Claude Sonnet 5.5
System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…
Benchmark / dataset
Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark
Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…