74 artifacts matching
Introducing ChatGPT for Teens: Built for learning, backed by protections
OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
System Card: Claude Sonnet 5
Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…
Sensing but not alerting: ChatGPT mental health triage gaps in simulated psychodermatology conversations
Research letter testing whether ChatGPT recognises mental health concerns and recommends appropriate referral during simulated multi-turn psychodermatology conversations. Fifty first-person narrative…
Development of a Consensus Statement to Guide AI Chatbot Responses to Suicide Risk Disclosure
An online modified Delphi study producing the first consensus statement on how generative AI chatbots should respond when a user discloses suicide risk, together with taxonomies of the potential harm…
PIPEDA Findings #2026-004: Commissioner-Initiated Complaint Concerning X Corp. and X.AI LLC
A commissioner-initiated federal investigation, conducted jointly with provincial privacy counterparts, into X Corp. and X.AI LLC's compliance with Canada's PIPEDA in connection with Grok's image-gen…
How AI Companies are Handling Suicide and Self-Harm Today
Drawing on a March 2026 multistakeholder workshop convening frontier AI companies, clinicians, researchers, and people with lived experience, Partnership on AI presents a taxonomy of six intervention…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
A benchmark dataset of 2,123 real-world Replika conversations annotated across nine safety risk categories (including sexual behavior, aggression, substance abuse, and manipulation) for evaluating LL…
Artificial intelligence-associated delusions and large language models: risks, mechanisms of delusion co-creation, and safeguarding strategies
A Personal View in The Lancet Psychiatry from a King's College London-led group examining how large language models may validate or amplify delusional or grandiose content in users vulnerable to psyc…
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
A taxonomy of 37 dark patterns in AI chatbots, organized into four high-level categories, spanning general-purpose systems (ChatGPT, Gemini, Claude) and companion platforms (Replika, Character.AI). S…