Skip to main content

Browse the library

The complete record — 223 artifacts, last updated 20 Aug 2026. Also available as JSON and RSS (CC BY 4.0).

Filters:

74 artifacts matching

18 Aug 2026 OpenAI Lab publication

Introducing ChatGPT for Teens: Built for learning, backed by protections

OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…

17 Aug 2026 Journal of Psychopathology and Clinical Science (American Psychological Association) Peer-reviewed

A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)

The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…

12 Aug 2026 xAI Lab publication

Model Card: Grok 4.6

36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…

6 Aug 2026 OpenAI Lab publication

GPT-5.6 – August Updates

System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…

6 Aug 2026 PsyArXiv (Universidad Francisco de Vitoria; Durham University) Preprint

Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns

Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…

2 Aug 2026 arXiv (Virginia Tech) Preprint

Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy

Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…

28 Jul 2026 Mistral AI Lab publication

Shieldstral

Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…

16 Jul 2026 Meta Lab publication

Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI

Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…

9 Jul 2026 Mila (Quebec AI Institute) & ROOST Lab publication

Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection

Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…

30 Jun 2026 Anthropic Lab publication

System Card: Claude Sonnet 5

Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…

25 Jun 2026 JAAD International (Elsevier, for the American Academy of Dermatology) Peer-reviewed

Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations

Research letter testing whether ChatGPT recognises mental health concerns and recommends appropriate referral during simulated multi-turn psychodermatology conversations. Fifty first-person narrative…

21 Jun 2026 PsyArXiv (Corporal Michael J. Crescenz VA Medical Center; University of Pennsylvania; Stanford; Columbia University and others) Preprint

Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure

An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…

11 Jun 2026 Office of the Privacy Commissioner of Canada (OPC) Regulator study

PIPEDA Findings #2026-004: Commissioner-Initiated Complaint Concerning X Corp. and X.AI LLC

A commissioner-initiated federal investigation, conducted jointly with provincial privacy counterparts, into X Corp. and X.AI LLC's compliance with Canada's PIPEDA in connection with Grok's image-gen…

11 Jun 2026 Partnership on AI NGO report

How AI Companies are Handling Suicide and Self-Harm Today

Drawing on a March 2026 multistakeholder workshop convening frontier AI companies, clinicians, researchers, and people with lived experience, Partnership on AI presents a taxonomy of six intervention…

9 Jun 2026 Anthropic Lab publication

System Card: Claude Fable 5 & Claude Mythos 5

Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…

3 Jun 2026 arXiv Benchmark / dataset

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…

1 Jun 2026 The Lancet Psychiatry Peer-reviewed

Artificial intelligence-associated delusions and large language models: risks, mechanisms of delusion co-creation, and safeguarding strategies

A Personal View in The Lancet Psychiatry from a King's College London-led group examining how large language models may validate or amplify delusional or grandiose content in users vulnerable to psyc…

29 May 2026 Center for Democracy & Technology (CDT Research) NGO report

Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design

A taxonomy of 37 dark patterns in AI chatbots, organized into four high-level categories, spanning general-purpose systems (ChatGPT, Gemini, Claude) and companion platforms (Replika, Character.AI). S…