Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

109 artifacts matching

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 Frontiers in Psychology (Frontiers Media) Peer-reviewed

Peer-reviewed

How emotional response styles of generative AI companions shape adaptive emotion regulation and short-term AI reliance

Two-condition randomized experiment with 380 adults in Guangzhou, China, comparing a generative-AI companion whose replies to negative emotion were adaptive regulation-oriented (validating the feelin…

7 Sept 2026 arXiv (University of Edinburgh; Middle East Technical University); accepted to EMNLP 2026 Preprint

Preprint

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…

4 Sept 2026 arXiv (Nanyang Technological University, School of Social Sciences) Preprint

Preprint

Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses

Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…

3 Sept 2026 arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics) Preprint

Preprint

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…

3 Sept 2026 arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026 Preprint

Preprint

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…

3 Sept 2026 Journal of Personality and Social Psychology (American Psychological Association) Peer-reviewed

Peer-reviewed

Judged by humans, comforted by machines: How psychological expectations shape experiences and perceptions of AI support

Examines how people experience and perceive social support from AI after minor aversive experiences such as rude interactions or social slights. Drawing on mind-perception theory, the authors propose…

3 Sept 2026 OpenAI Lab publication

Lab publication

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

1 Sept 2026 JMIR Mental Health (JMIR Publications) Peer-reviewed

Peer-reviewed

Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation

Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…

31 Aug 2026 Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) Peer-reviewed

Peer-reviewed

Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions

Pilot study assessing the clinical and cultural competence of a custom GPT configured to deliver systemic (couple and family therapy) interventions aligned with the AAMFT Code of Ethics and core comp…

29 Aug 2026 AI & SOCIETY (Springer) Peer-reviewed

Peer-reviewed

Benevolent Gravity: the lethal structure inherent in conversational AI design principles

Analytical paper arguing that the two dominant explanations for fatalities linked to conversational AI — safety-filter failure and commodified intimacy — are structurally insufficient, because a subs…

29 Aug 2026 arXiv (School of Computing and Information, University of Pittsburgh) Preprint

Preprint

Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

A perspectivist annotation study asking whether large language models used for distress detection capture the perspectives of the communities whose language they assess. 321 participants provided 9,5…

29 Aug 2026 ACM (Proceedings of Mensch und Computer 2026) Peer-reviewed

Peer-reviewed

Beyond Manipulation: How Users Perceive Harmful AI Chatbot Interactions

Mixed-methods study (N = 100) in which participants recalled a positive, an inappropriate or a manipulative chatbot interaction. Exploratory factor analysis of an adapted perceived-manipulation quest…

28 Aug 2026 Anthropic Lab publication

Lab publication

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…

27 Aug 2026 arXiv (University of Illinois Chicago; National University of Singapore) Preprint

Preprint

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…

26 Aug 2026 AI Security Institute (UK Department for Science, Innovation and Technology) Preprint

Preprint

When Do LLM Preferences Predict Downstream Behavior?

Study by the UK AI Security Institute testing whether large language models' elicited preferences predict their downstream behavior. Five frontier LLMs were evaluated across three domains: donation a…

26 Aug 2026 arXiv (Nanyang Technological University; National University of Singapore) Benchmark / dataset

Benchmark / dataset

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…

25 Aug 2026 arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) Benchmark / dataset

Benchmark / dataset

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…

21 Aug 2026 arXiv (preprint) Preprint

Preprint

Affective Context Amplifies Sycophancy in LLM Responses

A study of how a user's disclosed emotional state modulates sycophancy in subjective, evaluative exchanges. Drawing on ingratiation theory, the authors measure sycophancy as the divergence between a…