42 artifacts matching
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script s…
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
Preregistered experiment in which 1,500 UK adults each held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical across conditions and only the disc…
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
Pre-registered four-week longitudinal study (N=72, 182,451 lines of conversation) in which participants conversed with ChatGPT-4o either under a relational system prompt or unmodified, analysed throu…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Adolescent AI Use and Evaluations: Prospective Bidirectional Associations with Internalizing Symptoms
Two-wave longitudinal cohort study of 2,342 students in grades 6-8 across 22 middle schools in eight districts of a large southeastern US state, measuring AI chatbot use, use frequency and evaluation…
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Proposes a longitudinal evaluation framework (Theater-Stage-Judge) that uses persona-driven user simulation with dynamic psychological-state updating to assess the cognitive-developmental risks of AI…
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
A preprint examining how large language models handle psychological distress when it is entangled with delusional beliefs, using matched multi-turn simulations across clinically grounded personas and…
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…
What's on Your Mind? Exploring Privacy of Mental Health Apps
An empirical privacy analysis of 25 popular Android mental-health, therapy, and companion apps — including conversational-AI chatbots such as Replika, Talkie, Pi, Woebot, Wysa, and Youper — combining…