196 artifacts
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Interaction with AI companions and psychological well-being
Stanford-led study of 1,131 adult Character.AI users combining survey self-report with donated chat transcripts from 244 of them, analyzed with LLM-assisted methods against the Comprehensive Inventor…
Talking to machines: Children's experiences with AI assistants and companions
National survey report from Australia's eSafety Commissioner on children's use of AI assistants and companions, based on 1,950 children aged 10-17 surveyed in February-March 2026. Covers prevalence a…
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Introduces ANCHOR, an audit framework for long-horizon consistency in AI companions, evaluating persona enactment and trajectory recall over 2,008 conversations across 27 personas and four models. Fi…
Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…
Youth Perspectives on Online Safety, 2025
Sixth annual installment of Thorn's youth monitoring survey on US minors' online experiences and safety behaviors, surveying 1,003 minors aged 9-17 in late 2025. This wave adds questions on AI chatbo…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
Framework preprint examining how conversational AI systems affect users' psychological health, identifying benefits (information access, learning support) alongside risks including emotional entangle…
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Teens and Explicit Deepfakes in the Age of AI
Nationally representative survey of 1,314 US teens aged 13-17, fielded fall 2025, on exposure to and creation of AI-generated explicit sexual material. Documents widespread exposure, personal victimi…
Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of the AI Act
Commission guidelines (C(2026) 5054 final, adopted 20 July 2026) interpreting the AI Act Article 50 transparency obligations that apply from 2 August 2026, including the duty to disclose to users tha…
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study
Preregistered study comparing 49 large language models against 8 clinicians on detecting suicidal ideation embedded in psychotherapy transcripts of increasing length (0-200 speaker turns). Model F1 d…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Peer-reviewed benchmark and taxonomy for mental-health safety in LLMs, published in Findings of ACL 2026. R-MHSafe is a role-aware safety taxonomy characterizing clinically significant harm by the in…
Preliminary Report of the Independent International Scientific Panel on AI: Evidence-based assessment of opportunities, risks and impacts of AI
First report of the UN General Assembly-mandated Independent International Scientific Panel on AI, an independent body of scientists and experts from all five UN regions co-chaired by Yoshua Bengio a…
System Card: Claude Sonnet 5
Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…