92 artifacts matching
Preprint
Mitigating Social Sycophancy via Pluralistic Preference Optimization
Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…
Peer-reviewed
Artificial Intelligence–Associated Psychosis
A psychiatric case report from the mental health division of Northern Hospital, Melbourne, describes a 27-year-old woman with a year of psychotic symptoms whose presentation centred on ChatGPT. She u…
Benchmark / dataset
Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard
Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…
Peer-reviewed
Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…
Benchmark / dataset
FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy
Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…
Preprint
Sycophancy and Pressure Resistance in Medical Large Language Models: Systematic Review, Taxonomy, and Minimum Evaluation Framework
Registered systematic review of 31 benchmark and simulation studies on how medical LLMs change clinically relevant outputs when patient or clinician users introduce false premises, misleading evidenc…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Preprint
Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…
Preprint
Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversati…
Preprint
Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…
Lab publication
Model Card: Grok 4.7
Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…
Preprint
Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
A preregistered audit of six deployed assistants (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemma 4 31B IT, Mistral Small 3.2 24B, DeepSeek V4 Flash) using 7,500 scripted multi-turn conversations that rand…
Peer-reviewed
Workers shift their views and pay more when AI chatbots pander to their values
An exploratory study and two pre-registered experiments grounded in moral foundations theory test whether large language models can systematically influence users' decisions by framing recommendation…
Preprint
Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial
Turn-level audit of a GPT-4o career-reflection agent used in a randomised trial that had found agent participants ending less committed to their career plans and more doubtful than participants doing…
Preprint
The Adaptation Dilemma: Cultural Fit Does Not Guarantee Safety in Mental-Health LLMs
Conceptual paper arguing that cultural fit and safety are distinct properties of mental-health conversations with general-purpose chatbots. It sorts interaction harms into two classes, imposition (th…
Framework
AI Safety Governance Framework 3.0 (人工智能安全治理框架3.0)
Third edition of China's national AI Safety Governance Framework, released by TC260 at the opening of the 2026 National Cybersecurity Publicity Week under the guidance of the Cyberspace Administratio…
Preprint
Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble
Fine-tuning on synthetic third-person stories about humans, containing no AI characters at all, transfers those characters' conditional behaviours and implicit preferences into the assistant's conduc…
Preprint
Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions
Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…
Preprint
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…
Preprint
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…