66 artifacts matching
Preprint
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
Introduces SPINE, a benchmark in which a language-model proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns on 100 false-presupposition and 100 unet…
Preprint
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…
Preprint
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…
Benchmark / dataset
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…
Preprint
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with…
Preprint
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Extends personalized safety to vision-language models. MPS-Bench pairs 5,181 image-plus-query scenarios, built from 584 real-world images across 12 high-risk domains (including Health, Relationship,…
Lab publication
2026 Responsible AI Transparency Report
Microsoft's third annual responsible-AI transparency report, covering July 2025 to June 2026. One of its four 2026 trends is people turning to conversational AI for personal advice, health questions…
Benchmark / dataset
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Releases CompanionSim, a simulation framework and corpus of 2,240 simulated multi-turn human-chatbot conversations covering 16 chatbot behaviours across seven use cases, built to study AI companionsh…
Benchmark / dataset
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…
Lab publication
Funding better evaluations of AI's impact on wellbeing
Announcement of a $5 million grant programme funding independent research into how AI affects user wellbeing, offering money, model access and technical support to teams building open-source evaluati…
Benchmark / dataset
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…
Preprint
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
The authors introduce an ontology of ten therapeutic moves, compact function-based categories grounded in the MULTI-60 psychotherapy process inventory, validated through an annotation campaign with f…
Benchmark / dataset
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
Peer-reviewed
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Peer-reviewed Nature Medicine study introducing SIM-VAIL (simulated vulnerability-amplifying interaction loops), a clinically validated framework for auditing chatbot behavior in mental-health contex…
Preprint
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
Preprint
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Preprint
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…
Lab publication
Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…
Peer-reviewed
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…
Peer-reviewed
Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
Benchmark evaluation of how four consumer AI systems maintain safety boundaries when answering paediatric health questions from caregivers, using PediatricSafetyBench-v2: 300 authentic caregiver quer…