Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

42 artifacts matching

13 Aug 2026 arXiv preprint Preprint

How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior

Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script s…

12 Aug 2026 arXiv (National Taiwan University; University of Bamberg) Preprint

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion

Preregistered experiment in which 1,500 UK adults each held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical across conditions and only the disc…

11 Aug 2026 arXiv preprint (University of Warwick / Forensic Capability Network) Benchmark / dataset

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…

11 Aug 2026 arXiv preprint Preprint

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement

Pre-registered four-week longitudinal study (N=72, 182,451 lines of conversation) in which participants conversed with ChatGPT-4o either under a relational system prompt or unmodified, analysed throu…

8 Aug 2026 arXiv preprint Preprint

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…

6 Aug 2026 arXiv preprint Preprint

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…

6 Aug 2026 PsyArXiv (Universidad Francisco de Vitoria; Durham University) Preprint

Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns

Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…

5 Aug 2026 arXiv (Stanford-led author team) Preprint

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…

3 Aug 2026 PsyArXiv (University of North Carolina at Chapel Hill) Preprint

Adolescent AI Use and Evaluations: Prospective Bidirectional Associations with Internalizing Symptoms

Two-wave longitudinal cohort study of 2,342 students in grades 6-8 across 22 middle schools in eight districts of a large southeastern US state, measuring AI chatbot use, use frequency and evaluation…

2 Aug 2026 arXiv (Virginia Tech) Preprint

Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy

Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…

30 Jul 2026 arXiv (Salesforce AI Research) Preprint

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…

27 Jul 2026 arXiv preprint Preprint

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…

24 Jul 2026 arXiv preprint Preprint

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…

24 Jun 2026 arXiv (Shanghai AI Laboratory-led) Preprint

Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

Proposes a longitudinal evaluation framework (Theater-Stage-Judge) that uses persona-driven user simulation with dynamic psychological-state updating to assess the cognitive-developmental risks of AI…

21 Jun 2026 PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) Preprint

Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure

An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…

9 Jun 2026 arXiv (Emory University-led) Benchmark / dataset

Expert-Level Crisis Detection in Mental Health Conversations

A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…

3 Jun 2026 arXiv Benchmark / dataset

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…

31 May 2026 arXiv preprint Preprint

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

A preprint examining how large language models handle psychological distress when it is entangled with delusional beliefs, using matched multi-turn simulations across clinically grounded personas and…

27 May 2026 arXiv Benchmark / dataset

SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats

A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…

3 May 2026 arXiv (University of California, Irvine / University of California, Riverside) Preprint

What's on Your Mind? Exploring Privacy of Mental Health Apps

An empirical privacy analysis of 25 popular Android mental-health, therapy, and companion apps — including conversational-AI chatbots such as Replika, Talkie, Pi, Woebot, Wysa, and Youper — combining…