Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

191 artifacts matching

5 Oct 2026 PensionBee Industry survey

Industry survey

One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)

PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three…

1 Oct 2026 arXiv (Stanford University; University of Washington; Amazon) Preprint

Preprint

Mitigating Social Sycophancy via Pluralistic Preference Optimization

Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology Peer-reviewed

Peer-reviewed

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusi…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

30 Sept 2026 arXiv (Technical University of Munich; Massachusetts Institute of Technology); accepted at HICSS-60 (2027) Preprint

Preprint

Persona and Persuasive Framing in AI Voice Agents: A 2x2 Field Experiment with Children

Randomised 2x2 field experiment embedded in a public German-language Santa Claus telephone hotline before Christmas 2025, in which children's calls were routed to LLM voice agents that varied persona…

30 Sept 2026 Keurmerk Verantwoorde Affiliates (KVA), an initiative of XY Legal Solutions B.V. Industry survey

Industry survey

Generatieve AI en illegaal online gokaanbod: Hoe AI-tools Nederlandse consumenten bij niet-vergunde casino's brengen

Dutch-language test of ten consumer generative AI tools on whether simple, realistic questions lead users to online casinos without a Dutch licence. Each tool received nine prompts in three series: n…

28 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Sonnet 5.5

System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…

23 Sept 2026 ACM (Proceedings of the ACM on Human-Computer Interaction, vol. 10 no. 6, CSCW 2026) Peer-reviewed

Peer-reviewed

One Question, Four Voices: How Advice for Alzheimer's Caregiving Differs Between Caregivers, Clinicians, and Large Language Models

A side-by-side comparison of responses to 85 real caregiver questions from the ALZConnected forum across peer caregivers, physicians, ChatGPT and CareGPT, a retrieval-augmented variant grounded in pe…

23 Sept 2026 Proceedings of the ACM on Human-Computer Interaction (CSCW 2026); Brown University Peer-reviewed

Peer-reviewed

Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions

Compares session-level behaviour of human peer counsellors and an LLM counsellor trained on the same single-session CBT manual. An 18-month ethnography on a peer-support platform (seven counsellors,…

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

23 Sept 2026 PEC Innovation (Elsevier) Peer-reviewed

Peer-reviewed

Language and cultural prompts influence large language models' responses to culturally anchored mental health attitudes

Two factorial experiments (2 prompt languages, Chinese or English, by 2 cultural framings, China or the United States) presenting ChatGPT-4o and DeepSeek-V3 with the same description of a friend show…

22 Sept 2026 Anthropic Lab publication

Lab publication

Claude Opus 5.5 System Card

230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…

22 Sept 2026 arXiv (Northeastern University; University of Cambridge, Centre for Human-Inspired AI) Preprint

Preprint

Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus

Preregistered three-phase experiment (N=200 US adults) in which participants advised people facing real dilemmas and completed the PVQ-RR values questionnaire before, immediately after and one task a…

22 Sept 2026 arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University) Preprint

Preprint

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…

21 Sept 2026 arXiv (Cornell Tech) Preprint

Preprint

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…

21 Sept 2026 Journal of Broadcasting & Electronic Media (Taylor & Francis for the Broadcast Education Association) Peer-reviewed

Peer-reviewed

Authority Over Sympathy: How AI Emotional Tone and Thinking Dispositions Shape Hallucination Detection in Health Communication

Wizard-of-Oz experiment (N=476) testing how the emotional tone of AI-generated health responses that contain embedded hallucinations affects whether people detect the false content, framed by the Mis…

21 Sept 2026 xAI (styled SpaceXAI in the card) Lab publication

Lab publication

Model Card: Grok 4.7

Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…

20 Sept 2026 arXiv (Zhejiang University; National University of Singapore; Dartmouth College, Tuck School of Business; Singapore Sapiens Technology) Preprint

Preprint

AI Persona, Service Consumption, and User Intent Entropy: Field Experimental Evidence from an LLM Platform

A randomized field experiment on Agnes AI, a Singapore-based consumer chatbot platform, in which 9,586 newly registered users were assigned to a relational persona (warmer, more empathetic, more enga…

19 Sept 2026 arXiv (Purdue University, Department of Political Science) Preprint

Preprint

Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity

A preregistered audit of six deployed assistants (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemma 4 31B IT, Mistral Small 3.2 24B, DeepSeek V4 Flash) using 7,500 scripted multi-turn conversations that rand…