131 artifacts matching
Preprint
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…
Preprint
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…
Preprint
Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
Audits 31 pre-specified NLP techniques from seven methodological families (model scaling, synthetic data, loss reweighting, ensembling, structured prediction, threshold tuning and LLM methods) in rou…
Preprint
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
Conversation-analysis-grounded study of how language models respond when a user challenges their answer. Introduces a taxonomy of six challenge types and a four-layer response framework (whether the…
Preprint
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…
Peer-reviewed
Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data
Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…
Preprint
Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses
Factorial vignette experiment on how a language model's moral advice about eldercare changes under sustained user pushback. GPT-4o-mini received Chinese-language caregiving dilemmas in two framings (…
Benchmark / dataset
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…
Preprint
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Peer-reviewed
Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation
Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…
Benchmark / dataset
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…
Preprint
Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries
An audit of what three free consumer generative-search products (ChatGPT, Perplexity, Google AI Overview) cite when answering mental-health questions. Twenty English questions were run under two prom…
Peer-reviewed
Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis
PRISMA systematic review of quantitative studies evaluating general-purpose large language models for direct mental health care tasks, searched across PubMed, Embase, ACM Digital Library, IEEE Xplore…
Preprint
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts
A perspectivist annotation study asking whether large language models used for distress detection capture the perspectives of the communities whose language they assess. 321 participants provided 9,5…
Lab publication
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…
Preprint
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…
Preprint
When Do LLM Preferences Predict Downstream Behavior?
Study by the UK AI Security Institute testing whether large language models' elicited preferences predict their downstream behavior. Five frontier LLMs were evaluated across three domains: donation a…
Benchmark / dataset
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…
Benchmark / dataset
HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark ho…