Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

46 artifacts matching

8 Sept 2026 arXiv (Grow Therapy; Stanford University School of Medicine) Preprint

Preprint

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 arXiv (Stanford University) Preprint

Preprint

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…

7 Sept 2026 arXiv (University of Edinburgh; Middle East Technical University); accepted to EMNLP 2026 Preprint

Preprint

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…

7 Sept 2026 arXiv (King's College London Institute of Psychiatry, Psychology and Neuroscience; South London and Maudsley NHS Foundation Trust; The Human Line Project) Preprint

Preprint

Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports

Cross-sectional secondary analysis of 185 deidentified accounts of mental-health harm linked with AI chatbot use (95 first-hand, 90 from relatives, partners or friends) submitted through the web form…

6 Sept 2026 UK AI Security Institute (Societal Impacts Team), posted on PsyArXiv Preprint

Preprint

How is AI impacting people?

Narrative review by the UK AI Security Institute's Societal Impacts Team organised around eight questions the public most often raises about AI's effects on people, including whether AI is underminin…

4 Sept 2026 arXiv (Nanyang Technological University, School of Social Sciences) Preprint

Preprint

Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses

Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…

3 Sept 2026 arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics) Preprint

Preprint

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

28 Aug 2026 Anthropic Lab publication

Lab publication

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…

27 Aug 2026 arXiv (University of Illinois Chicago; National University of Singapore) Preprint

Preprint

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…

21 Aug 2026 arXiv (preprint) Preprint

Preprint

Affective Context Amplifies Sycophancy in LLM Responses

A study of how a user's disclosed emotional state modulates sycophancy in subjective, evaluative exchanges. Drawing on ingratiation theory, the authors measure sycophancy as the divergence between a…

20 Aug 2026 npj Digital Medicine (Nature Portfolio) Peer-reviewed

Peer-reviewed

A scoping review on the mental health harms of LLM-based chatbots

A PRISMA-based scoping review synthesising research on mental health harms associated with chatbots built on large language models. A systematic search with a validated search string across five data…

17 Aug 2026 Journal of Psychopathology and Clinical Science (American Psychological Association) Peer-reviewed

Peer-reviewed

A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)

The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…

12 Aug 2026 xAI Lab publication

Lab publication

Model Card: Grok 4.6

36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…

6 Aug 2026 arXiv preprint Preprint

Preprint

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…

2 Aug 2026 arXiv (Virginia Tech) Preprint

Preprint

Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy

Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…

1 Aug 2026 Joint Commission on Technology and Science (JCOTS), Virginia General Assembly, Division of Legislative Services Government report

Government report

Supplementary Report: AI Chatbots: Companionship & Minors

Staff report of Virginia's legislative technology commission, supplementing its 2025 AI-chatbot study after two bills, HB 635 (Artificial Intelligence Chatbots Act) and SB 796 (Artificial Intelligenc…

15 Jul 2026 arXiv (Yonsei University; CASA Labs; Fudan University; St. Johnsbury Academy Jeju) Preprint

Preprint

The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

Controlled experiment testing whether a user's expressed emotional distress shifts commercial LLMs toward endorsing premature, consequential life decisions (quitting a stable job, expanding a busines…

13 Jul 2026 Anthropic Lab publication

Lab publication

Claude's values across models and languages

An observational study of 309,815 anonymised production conversations characterising the values an assistant expresses and how that expression varies by model version and by the user's language. Expr…