Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

181 artifacts matching

1 Oct 2026 arXiv (Stanford University; University of Washington; Amazon) Preprint

Preprint

Mitigating Social Sycophancy via Pluralistic Preference Optimization

Proposes Pluralistic Preference Optimization (PlurPO), a post-training method in which a model simulates the stakeholders affected by a user's interpersonal situation and is trained to prefer respons…

30 Sept 2026 arXiv (ELLIS Institute Tübingen; Max Planck Institute for Intelligent Systems; Tübingen AI Center) Benchmark / dataset

Benchmark / dataset

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Benchmark that scores sycophancy and calibrated validation (acknowledging a user's feelings without yielding) as separate axes over ten-turn conversations driven by an adaptive user simulator. It rel…

30 Sept 2026 arXiv (Technical University of Munich; Massachusetts Institute of Technology); accepted at HICSS-60 (2027) Preprint

Preprint

Persona and Persuasive Framing in AI Voice Agents: A 2x2 Field Experiment with Children

Randomised 2x2 field experiment embedded in a public German-language Santa Claus telephone hotline before Christmas 2025, in which children's calls were routed to LLM voice agents that varied persona…

29 Sept 2026 arXiv (Rutgers University; independent researcher, Kolkata) Preprint

Preprint

How People Use ChatGPT: Conversation-Level Evidence from India, Nigeria, Brazil, and Pakistan

Data-donation study of complete ChatGPT exports from 1,252 users in India, Nigeria, Brazil and Pakistan (202,590 conversations, December 2022 to February 2026), paired with self-reported age and gend…

29 Sept 2026 arXiv (Massachusetts Institute of Technology); accepted to Findings of EMNLP 2026 Benchmark / dataset

Benchmark / dataset

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Benchmark comparing language-model and physician triage recommendations (self-manage at home, in-person visit, tests or referral) on clinical cases, including patient-written Reddit r/AskDocs posts,…

28 Sept 2026 arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University) Benchmark / dataset

Benchmark / dataset

Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark

Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross…

28 Sept 2026 arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin) Preprint

Preprint

Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT

Mixed-methods study of how young adults aged 18 to 25 use ChatGPT when distressed, built on complete donated ChatGPT histories (19,930 conversations from 158 participants) plus a survey that included…

27 Sept 2026 arXiv (Carnegie Mellon University; University Hospitals Cleveland Medical Center; Case Western Reserve University) Preprint

Preprint

Medical Knowledge Is Not All You Need: When Medical Q&A Becomes Situated Patient Assistance

Study of the questions skin cancer patients asked a voice assistant while practising postoperative wound care, followed by an offline replay of those questions to general-purpose language models. Man…

22 Sept 2026 arXiv (Northeastern University; University of Cambridge, Centre for Human-Inspired AI) Preprint

Preprint

Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus

Preregistered three-phase experiment (N=200 US adults) in which participants advised people facing real dilemmas and completed the PVQ-RR values questionnaire before, immediately after and one task a…

22 Sept 2026 arXiv (Harvard Kennedy School; Harvard Department of Statistics; Stanford University) Preprint

Preprint

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

Argues that social-sycophancy evaluations conflate inappropriate deference with conversational receptiveness, a social-psychology construct for engaging with a view one does not share. Using the Redd…

22 Sept 2026 arXiv (Carnegie Mellon University) Preprint

Preprint

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversati…

21 Sept 2026 arXiv (Northeastern University; University College Dublin) Preprint

Preprint

Annie, Are You Okay? How Style- and Context-Based Personalization Shape AI-Assisted Decision-Making

Preregistered 2 x 2 between-subjects experiment (N=240) in which participants ranked three comparably viable stocks, discussed them with a GPT-5.4-based assistant whose personalisation was varied by…

21 Sept 2026 arXiv (Cornell Tech) Preprint

Preprint

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) t…

21 Sept 2026 arXiv (Foundation AI, Cisco; Carnegie Mellon University) Preprint

Preprint

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Tests whether personal AI agents given a user's email inbox and profile steer economic recommendations by inferred wealth. Across roughly 325,000 runs on 13 models from four families in three decisio…

21 Sept 2026 arXiv (University of Massachusetts Amherst; University of Illinois Urbana-Champaign; Indiana University Indianapolis) Preprint

Preprint

Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors

Scenario-driven study of how three LLMs (ChatGPT 5.4 Thinking, Claude Sonnet 4.6 Extended Thinking, Gemini 3 Thinking) identify and resolve therapeutic-alliance ruptures across 21 mental-health conve…

20 Sept 2026 arXiv (Zhejiang University; National University of Singapore; Dartmouth College, Tuck School of Business; Singapore Sapiens Technology) Preprint

Preprint

AI Persona, Service Consumption, and User Intent Entropy: Field Experimental Evidence from an LLM Platform

A randomized field experiment on Agnes AI, a Singapore-based consumer chatbot platform, in which 9,586 newly registered users were assigned to a relational persona (warmer, more empathetic, more enga…

20 Sept 2026 PsyArXiv (OSF); Boston University; Purdue University; Brown University; Huainan Union University; Wuhan University; Harvard University Preprint

Preprint

Everyday use of AI for emotion regulation depends on the person and the situation

Thirteen-day ecological momentary assessment study in which 3,048 Chinese university students aged 18 to 29 reported emotional situations, emotion-regulation strategies and affect four times a day, y…

19 Sept 2026 arXiv (Purdue University, Department of Political Science) Preprint

Preprint

Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity

A preregistered audit of six deployed assistants (Claude Opus 4.8, GPT-5.5, Grok 4.3, Gemma 4 31B IT, Mistral Small 3.2 24B, DeepSeek V4 Flash) using 7,500 scripted multi-turn conversations that rand…

18 Sept 2026 PsyArXiv (OSF); University of British Columbia (Psychiatry, Data Science Institute, Population and Public Health, Computer Science, Medicine) Preprint

Preprint

AI-based detection of suicidal ideation in text: model development and evaluation for a student mental health chatbot

Development and evaluation of a lightweight suicidal-ideation detection system intended for integration into Minder, a University of British Columbia mental-health chatbot for students. A fine-tuned…

17 Sept 2026 arXiv (Georgia Institute of Technology) Preprint

Preprint

Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation

Introduces COPES (Community-centered Peer Engaged Support), a dataset of mental-health support-seeking Reddit queries with community-endorsed responses (5,536 posts after filtering, across five commu…