Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

642 artifacts

5 Oct 2026 PensionBee Industry survey

Industry survey

One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)

PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three…

5 Oct 2026 SOS-Kinderdörfer weltweit; Terre des Hommes Deutschland NGO report

NGO report

Sicher mit KI aufwachsen. Zwischen Algorithmen, Chatbots und Deepfakes: Schutz, Teilhabe und Bildung für Kinder und Jugendliche (Vorstudie)

A 67-page German-language pre-study commissioned by the child-rights organisations SOS-Kinderdörfer weltweit and Terre des Hommes and written with Klartext AI. It maps where children encounter AI, th…

4 Oct 2026 CLEF 2026 Working Notes (CEUR Workshop Proceedings Vol-4283); Universidade da Coruña IRLab, University of Sheffield, Università della Svizzera italiana Benchmark / dataset

Benchmark / dataset

Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)

Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas…

1 Oct 2026 arXiv (Stanford University; University of Washington; Amazon) Preprint

Preprint

Mitigating Social Sycophancy via Pluralistic Preference Optimization

Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…

1 Oct 2026 Legal Ombudsman (England and Wales) Government report

Government report

AI and complaints: removing barriers, reinforcing divides? How AI is influencing who complains, how they complain, and what could come next

Research briefing from the Legal Ombudsman for England and Wales on how consumers use generative AI when deciding whether and how to complain about regulated services. It combines an Ipsos survey of…

1 Oct 2026 The Primary Care Companion for CNS Disorders (Physicians Postgraduate Press); Northern Hospital, Melbourne; PGIMER Chandigarh Peer-reviewed

Peer-reviewed

Artificial Intelligence–Associated Psychosis

A psychiatric case report from the mental health division of Northern Hospital, Melbourne, describes a 27-year-old woman with a year of psychotic symptoms whose presentation centred on ChatGPT. She u…

1 Oct 2026 Research Square (preprint); Spring Health (Spring Care Inc) Preprint

Preprint

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology Peer-reviewed

Peer-reviewed

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

30 Sept 2026 arXiv (Technical University of Munich; Massachusetts Institute of Technology); accepted at HICSS-60 (2027) Preprint

Preprint

Persona and Persuasive Framing in AI Voice Agents: A 2x2 Field Experiment with Children

Randomised 2x2 field experiment embedded in a public German-language Santa Claus telephone hotline before Christmas 2025, in which children's calls were routed to LLM voice agents that varied persona…

30 Sept 2026 Keurmerk Verantwoorde Affiliates (KVA), an initiative of XY Legal Solutions B.V. Industry survey

Industry survey

Generatieve AI en illegaal online gokaanbod: Hoe AI-tools Nederlandse consumenten bij niet-vergunde casino's brengen

Dutch-language test of ten consumer generative AI tools on whether simple, realistic questions lead users to online casinos without a Dutch licence. Each tool received nine prompts in three series: n…

30 Sept 2026 Sapien Labs (Global Mind Project) NGO report

NGO report

Generative AI Use and Mind Health Outcomes

Sapien Labs rapid report analysing generative AI chatbot use among 264,085 adults in the Global Mind Project's online survey across 85+ countries, relating frequency and purpose of use to the Mind He…

30 Sept 2026 Suicide Policy Research (Japan Suicide Countermeasures Promotion Center); Specified Nonprofit Corporation OVA; Kyoto University; Wako University Peer-reviewed

Peer-reviewed

AI or Human Support for Suicide Prevention? Examining Help-Seeking Intention in Suicidal Crisis

A cross-sectional web survey of 1,024 Japanese adults aged 18-69 with severe psychological distress (Kessler-6 score of 13 or more) asks whether they would use chat-based crisis support delivered by…

29 Sept 2026 arXiv (Rutgers University; independent researcher, Kolkata) Preprint

Preprint

How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan

Data-donation study of complete ChatGPT exports from 1,252 users in India, Nigeria, Brazil and Pakistan (202,590 conversations, December 2022 to February 2026), paired with self-reported age and gend…

29 Sept 2026 arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026 Benchmark / dataset

Benchmark / dataset

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…

29 Sept 2026 Cyberpsychology, Behavior, and Social Networking (SAGE); Tel-Hai College; Bar-Ilan University; University of Haifa Peer-reviewed

Peer-reviewed

AI-Mediated Mental Health Support: The Role of Attachment Orientation and Psychological Distress

A preregistered cross-sectional survey of 584 Israeli adults who use general-purpose generative AI asks whether attachment orientation (ECR-RS) and psychological distress (DASS-21) are associated wit…

29 Sept 2026 medRxiv (preprint); University Hospital Frankfurt; UKP Lab, Technical University of Darmstadt Preprint

Preprint

From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue

The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…

28 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Sonnet 5.5

System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…

28 Sept 2026 arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University) Benchmark / dataset

Benchmark / dataset

Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark

Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…