Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

92 artifacts matching

1 Oct 2026 arXiv (Stanford University; University of Washington; Amazon) Preprint

Preprint

Mitigating Social Sycophancy via Pluralistic Preference Optimization

Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…

1 Oct 2026 The Primary Care Companion for CNS Disorders (Physicians Postgraduate Press); Northern Hospital, Melbourne; PGIMER Chandigarh Peer-reviewed

Peer-reviewed

Artificial Intelligence–Associated Psychosis

A psychiatric case report from the mental health division of Northern Hospital, Melbourne, describes a 27-year-old woman with a year of psychotic symptoms whose presentation centred on ChatGPT. She u…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology Peer-reviewed

Peer-reviewed

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

23 Sept 2026 JMIR Preprints (JMIR Publications) Preprint

Preprint

Sycophancy and Pressure Resistance in Medical Large Language Models: Systematic Review, Taxonomy, and Minimum Evaluation Framework

Registered systematic review of 31 benchmark and simulation studies on how medical LLMs change clinically relevant outputs when patient or clinician users introduce false premises, misleading evidenc…

22 Sept 2026 Anthropic Lab publication

Lab publication

Claude Opus 5.5 System Card

230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…

22 Sept 2026 arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University) Preprint

Preprint

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…

22 Sept 2026 arXiv (Carnegie Mellon University) Preprint

Preprint

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversati…

21 Sept 2026 arXiv (Cornell Tech) Preprint

Preprint

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…

21 Sept 2026 xAI (styled SpaceXAI in the card) Lab publication

Lab publication

Model Card: Grok 4.7

Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…

19 Sept 2026 arXiv (Purdue University, Department of Political Science) Preprint

Preprint

Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity

A preregistered audit of six deployed assistants (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemma 4 31B IT, Mistral Small 3.2 24B, DeepSeek V4 Flash) using 7,500 scripted multi-turn conversations that rand…

18 Sept 2026 Scientific Reports (Springer Nature); Australian National University; University of Cambridge; University of Chicago Peer-reviewed

Peer-reviewed

Workers shift their views and pay more when AI chatbots pander to their values

An exploratory study and two pre-registered experiments grounded in moral foundations theory test whether large language models can systematically influence users' decisions by framing recommendation…

17 Sept 2026 arXiv (Stanford University; University of Virginia; NAVER Cloud) Preprint

Preprint

Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial

Turn-level audit of a GPT-4o career-reflection agent used in a randomised trial that had found agent participants ending less committed to their career plans and more doubtful than participants doing…

16 Sept 2026 PsyArXiv (OSF); McGill University (Department of Psychiatry; Department of Philosophy); The Decision Lab Preprint

Preprint

The Adaptation Dilemma: Cultural Fit Does Not Guarantee Safety in Mental-Health LLMs

Conceptual paper arguing that cultural fit and safety are distinct properties of mental-health conversations with general-purpose chatbots. It sorts interaction harms into two classes, imposition (th…

14 Sept 2026 National Technical Committee 260 on Cybersecurity of SAC (TC260), issued under the guidance of the Cyberspace Administration of China (CAC) Framework

Framework

AI Safety Governance Framework 3.0 (人工智能安全治理框架3.0)

Third edition of China's national AI Safety Governance Framework, released by TC260 at the opening of the 2026 National Cybersecurity Publicity Week under the guidance of the Cyberspace Administratio…

9 Sept 2026 arXiv (Truthful AI; Harvard University; METR; University of Oxford) Preprint

Preprint

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

Fine-tuning on synthetic third-person stories about humans, containing no AI characters at all, transfers those characters' conditional behaviours and implicit preferences into the assistant's conduc…

8 Sept 2026 arXiv (Grow Therapy; Stanford University School of Medicine) Preprint

Preprint

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 arXiv (Stanford University) Preprint

Preprint

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…