109 artifacts matching
Preprint
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…
Peer-reviewed
How emotional response styles of generative AI companions shape adaptive emotion regulation and short-term AI reliance
Two-condition randomized experiment with 380 adults in Guangzhou, China, comparing a generative-AI companion whose replies to negative emotion were adaptive regulation-oriented (validating the feelin…
Preprint
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…
Preprint
Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses
Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…
Preprint
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…
Preprint
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…
Peer-reviewed
Judged by humans, comforted by machines: How psychological expectations shape experiences and perceptions of AI support
Examines how people experience and perceive social support from AI after minor aversive experiences such as rude interactions or social slights. Drawing on mind-perception theory, the authors propose…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Peer-reviewed
Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation
Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…
Peer-reviewed
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Pilot study assessing the clinical and cultural competence of a custom GPT configured to deliver systemic (couple and family therapy) interventions aligned with the AAMFT Code of Ethics and core comp…
Peer-reviewed
Benevolent Gravity: the lethal structure inherent in conversational AI design principles
Analytical paper arguing that the two dominant explanations for fatalities linked to conversational AI — safety-filter failure and commodified intimacy — are structurally insufficient, because a subs…
Preprint
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts
A perspectivist annotation study asking whether large language models used for distress detection capture the perspectives of the communities whose language they assess. 321 participants provided 9,5…
Peer-reviewed
Beyond Manipulation: How Users Perceive Harmful AI Chatbot Interactions
Mixed-methods study (N = 100) in which participants recalled a positive, an inappropriate or a manipulative chatbot interaction. Exploratory factor analysis of an adapted perceived-manipulation quest…
Lab publication
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…
Preprint
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…
Preprint
When Do LLM Preferences Predict Downstream Behavior?
Study by the UK AI Security Institute testing whether large language models' elicited preferences predict their downstream behavior. Five frontier LLMs were evaluated across three domains: donation a…
Benchmark / dataset
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…
Benchmark / dataset
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…
Preprint
Affective Context Amplifies Sycophancy in LLM Responses
A study of how a user's disclosed emotional state modulates sycophancy in subjective, evaluative exchanges. Drawing on ingratiation theory, the authors measure sycophancy as the divergence between a…