Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

18 artifacts matching

17 Sept 2026 Google DeepMind (arXiv preprint) Preprint

Preprint

Tailored to you: longitudinal effects of personalising language models

Five-day study in which 992 Prolific participants completed a daily advice-seeking conversation with a language model, randomised to a non-personalised control, a memory-based personalisation conditi…

8 Sept 2026 Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University Peer-reviewed

Peer-reviewed

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Peer-reviewed MLHC 2026 paper (version of record of the July 2026 arXiv preprint) testing whether large language models identify and correct false presuppositions in patients' questions as a conversa…

13 Aug 2026 arXiv preprint Preprint

Preprint

How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior

Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script s…

11 Aug 2026 arXiv preprint (University of Warwick / Forensic Capability Network) Benchmark / dataset

Benchmark / dataset

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…

11 Aug 2026 arXiv preprint Preprint

Preprint

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement

Pre-registered four-week longitudinal study (N=72, 182,451 lines of conversation) in which participants conversed with ChatGPT-4o either under a relational system prompt or unmodified, analysed throu…

8 Aug 2026 arXiv preprint Preprint

Preprint

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…

6 Aug 2026 arXiv preprint Preprint

Preprint

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…

27 Jul 2026 arXiv preprint Preprint

Preprint

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…

24 Jul 2026 arXiv preprint Preprint

Preprint

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…

1 Jul 2026 Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers) Peer-reviewed

Peer-reviewed

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

Version of record of the persona-grounded companion safety framework previously held as an arXiv preprint. Presents an end-to-end framework for controlled simulation and safety evaluation of multi-tu…

18 Jun 2026 OpenAI (Alignment Research Blog; arXiv preprint) Lab publication

Lab publication

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

OpenAI alignment research asking whether reinforcement learning on realistic conversations that reward beneficial traits (truthfulness, fairness, risk awareness, corrigibility, epistemic humility, co…

31 May 2026 arXiv preprint Preprint

Preprint

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

A preprint examining how large language models handle psychological distress when it is entangled with delusional beliefs, using matched multi-turn simulations across clinically grounded personas and…

14 May 2026 Meta (arXiv preprint) Lab publication

Lab publication

Muse Spark Safety & Preparedness Report

Meta's safety and preparedness report for Muse Spark, the model behind the Meta AI assistant. It presents catastrophic-risk evaluations under Meta's Advanced AI Scaling Framework, jailbreak and agent…

30 Apr 2026 arXiv preprint Preprint superseded

Preprint

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

Presents an end-to-end simulation framework for evaluating AI companion app safety across multi-turn conversations, using nine clinically-grounded vulnerable personas (including major depressive diso…

7 Apr 2026 arXiv preprint Benchmark / dataset superseded

Benchmark / dataset

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Introduces a taxonomy of elderly-specific risks in LLM chatbot interactions (3 levels, 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains) grounded…

11 Jan 2026 arXiv preprint Preprint

Preprint

Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse

An expert-led evaluation of four large language models — two general-purpose and two domain-specific for intimate partner violence contexts — responding to real-world questions about technology-facil…

13 May 2025 OpenAI (arXiv preprint) Benchmark / dataset

Benchmark / dataset

HealthBench: Evaluating Large Language Models Towards Improved Human Health

Open-source benchmark from OpenAI that measures the performance and safety of large language models in health conversations. It consists of 5,000 multi-turn conversations between a model and an indiv…

21 Mar 2025 MIT Media Lab; OpenAI (arXiv preprint) Preprint

Preprint

How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study

A four-week randomized controlled experiment in which 981 participants used ChatGPT (GPT-4o) for at least five minutes a day under one of nine conditions crossing interaction mode (text, neutral voic…