46 artifacts matching
Preprint
Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions
Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…
Preprint
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…
Preprint
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…
Preprint
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…
Preprint
Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
Cross-sectional secondary analysis of 185 deidentified accounts of mental-health harm linked with AI chatbot use (95 first-hand, 90 from relatives, partners or friends) submitted through the web form…
Preprint
How is AI impacting people?
Narrative review by the UK AI Security Institute's Societal Impacts Team organised around eight questions the public most often raises about AI's effects on people, including whether AI is underminin…
Preprint
Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses
Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…
Preprint
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Lab publication
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…
Preprint
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…
Preprint
Affective Context Amplifies Sycophancy in LLM Responses
A study of how a user's disclosed emotional state modulates sycophancy in subjective, evaluative exchanges. Drawing on ingratiation theory, the authors measure sycophancy as the divergence between a…
Peer-reviewed
A scoping review on the mental health harms of LLM-based chatbots
A PRISMA-based scoping review synthesising research on mental health harms associated with chatbots built on large language models. A systematic search with a validated search string across five data…
Peer-reviewed
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…
Lab publication
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
Preprint
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
Preprint
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Government report
Supplementary Report: AI Chatbots: Companionship & Minors
Staff report of Virginia's legislative technology commission, supplementing its 2025 AI-chatbot study after two bills, HB 635 (Artificial Intelligence Chatbots Act) and SB 796 (Artificial Intelligenc…
Preprint
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
Controlled experiment testing whether a user's expressed emotional distress shifts commercial LLMs toward endorsing premature, consequential life decisions (quitting a stable job, expanding a busines…
Lab publication
Claude's values across models and languages
An observational study of 309,815 anonymised production conversations characterising the values an assistant expresses and how that expression varies by model version and by the user's language. Expr…