11 artifacts matching
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script s…
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
Pre-registered four-week longitudinal study (N=72, 182,451 lines of conversation) in which participants conversed with ChatGPT-4o either under a relational system prompt or unmodified, analysed throu…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
A preprint examining how large language models handle psychological distress when it is entangled with delusional beliefs, using matched multi-turn simulations across clinically grounded personas and…
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Presents an end-to-end simulation framework for evaluating AI companion app safety across multi-turn conversations, using nine clinically-grounded vulnerable personas (including major depressive diso…
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
Introduces a taxonomy of elderly-specific risks in LLM chatbot interactions (3 levels, 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains) grounded…
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
An expert-led evaluation of four large language models — two general-purpose and two domain-specific for intimate partner violence contexts — responding to real-world questions about technology-facil…