Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

131 artifacts matching

8 Sept 2026 arXiv (Texas A&M University; University of Cincinnati) Preprint

Preprint

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…

8 Sept 2026 arXiv (Stanford University) Preprint

Preprint

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…

7 Sept 2026 arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County) Preprint

Preprint

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

Audits 31 pre-specified NLP techniques from seven methodological families (model scaling, synthetic data, loss reweighting, ensembling, structured prediction, threshold tuning and LLM methods) in rou…

7 Sept 2026 arXiv (University of Edinburgh; Middle East Technical University); accepted to EMNLP 2026 Preprint

Preprint

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…

5 Sept 2026 arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 Preprint

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…

5 Sept 2026 npj Digital Medicine (Springer Nature) Peer-reviewed

Peer-reviewed

Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data

Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…

4 Sept 2026 arXiv (Nanyang Technological University, School of Social Sciences) Preprint

Preprint

Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses

Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…

3 Sept 2026 arXiv (National University of Singapore, Department of Computer Science) Benchmark / dataset

Benchmark / dataset

LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues

Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…

3 Sept 2026 arXiv (University of California, Los Angeles; University of Washington; Microsoft Research Asia; William & Mary); accepted at COLM 2026 Preprint

Preprint

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…

3 Sept 2026 OpenAI Lab publication

Lab publication

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…

1 Sept 2026 JMIR Mental Health (JMIR Publications) Peer-reviewed

Peer-reviewed

Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation

Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…

31 Aug 2026 arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) Benchmark / dataset

Benchmark / dataset

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…

31 Aug 2026 arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School; Pontificia Universidad Javeriana; Tufts University School of Medicine) Preprint

Preprint

Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries

An audit of what three free consumer generative-search products (ChatGPT, Perplexity, Google AI Overview) cite when answering mental-health questions. Twenty English questions were run under two prom…

31 Aug 2026 JMIR AI (JMIR Publications) Peer-reviewed

Peer-reviewed

Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis

PRISMA systematic review of quantitative studies evaluating general-purpose large language models for direct mental health care tasks, searched across PubMed, Embase, ACM Digital Library, IEEE Xplore…

29 Aug 2026 arXiv (School of Computing and Information, University of Pittsburgh) Preprint

Preprint

Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

A perspectivist annotation study asking whether large language models used for distress detection capture the perspectives of the communities whose language they assess. 321 participants provided 9,5…

28 Aug 2026 Anthropic Lab publication

Lab publication

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…

27 Aug 2026 arXiv (University of Illinois Chicago; National University of Singapore) Preprint

Preprint

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…

26 Aug 2026 AI Security Institute (UK Department for Science, Innovation and Technology) Preprint

Preprint

When Do LLM Preferences Predict Downstream Behavior?

Study by the UK AI Security Institute testing whether large language models' elicited preferences predict their downstream behavior. Five frontier LLMs were evaluated across three domains: donation a…

26 Aug 2026 arXiv (Nanyang Technological University; National University of Singapore) Benchmark / dataset

Benchmark / dataset

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…

26 Aug 2026 arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo) Benchmark / dataset

Benchmark / dataset

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark ho…