Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

66 artifacts matching

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 arXiv (Stanford University) Preprint

Preprint

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…

5 Sept 2026 arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 Preprint

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…

3 Sept 2026 arXiv (National University of Singapore, Department of Computer Science) Benchmark / dataset

Benchmark / dataset

LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues

Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…

3 Sept 2026 arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics) Preprint

Preprint

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…

3 Sept 2026 arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026 Preprint

Preprint

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…

1 Sept 2026 Microsoft (Office of Responsible AI) Lab publication

Lab publication

2026 Responsible AI Transparency Report

Microsoft's third annual responsible-AI transparency report, covering July 2025 to June 2026. One of its four 2026 trends is people turning to conversational AI for personal advice, health questions…

31 Aug 2026 arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) Benchmark / dataset

Benchmark / dataset

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…

26 Aug 2026 arXiv (Nanyang Technological University; National University of Singapore) Benchmark / dataset

Benchmark / dataset

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…

25 Aug 2026 Anthropic Lab publication

Lab publication

Funding better evaluations of AI's impact on wellbeing

Announcement of a $5 million grant programme funding independent research into how AI affects user wellbeing, offering money, model access and technical support to teams building open-source evaluati…

25 Aug 2026 arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) Benchmark / dataset

Benchmark / dataset

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…

21 Aug 2026 arXiv (preprint) Preprint

Preprint

Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

The authors introduce an ontology of ten therapeutic moves, compact function-based categories grounded in the MULTI-60 psychotherapy process inventory, validated through an annotation campaign with f…

11 Aug 2026 arXiv preprint (University of Warwick / Forensic Capability Network) Benchmark / dataset

Benchmark / dataset

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…

7 Aug 2026 Nature Medicine Peer-reviewed

Peer-reviewed

A clinically validated framework for auditing AI chatbot behavior in mental health interactions

Peer-reviewed Nature Medicine study introducing SIM-VAIL (simulated vulnerability-amplifying interaction loops), a clinically validated framework for auditing chatbot behavior in mental-health contex…

6 Aug 2026 arXiv preprint Preprint

Preprint

Measuring and Detecting Harmful AI Sycophancy

Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…

5 Aug 2026 arXiv (Stanford-led author team) Preprint

Preprint

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…

30 Jul 2026 arXiv (Salesforce AI Research) Preprint

Preprint

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…

28 Jul 2026 Mistral AI Lab publication

Lab publication

Shieldstral

Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Peer-reviewed

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…

14 Jul 2026 npj Digital Medicine (Springer Nature) Peer-reviewed

Peer-reviewed

Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions

Benchmark evaluation of how four consumer AI systems maintain safety boundaries when answering paediatric health questions from caregivers, using PediatricSafetyBench-v2: 300 authentic caregiver quer…