27 artifacts matching
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detectio…
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta newsroom announcement that supervising parents using Instagram parental supervision will be proactively alerted when a teen's conversation with Meta AI suggests possible suicide or self-harm ris…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
Peer-reviewed study introducing a taxonomy of six clinically-informed mental-health crisis categories, an evaluation dataset of over 2,000 user inputs drawn from twelve public conversational datasets…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
Expert-Level Crisis Detection in Mental Health Conversations
A preprint introducing CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in multi-turn mental-health conversations, extending the same research group's earlier static-t…
System Card: Claude Opus 4.8
Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from pub…
Helping ChatGPT better recognize context in sensitive conversations
OpenAI post describing safety updates that let ChatGPT recognize risk emerging over the course of a conversation rather than judging each message alone, focused on suicide, self-harm and harm-to-othe…
Risks and Harms of Conversational Artificial Intelligence (CAI) Chatbot Use Among US Youth
A peer-reviewed, nationally representative survey of 3,466 US youth aged 13-17 on conversational AI chatbot use, motivations, and exposure to harmful chatbot behaviors. It quantifies adoption and the…
GPT-5.5 System Card
OpenAI's system card for GPT-5.5, published on its Deployment Safety Hub, documenting safety evaluations for the model. It includes a dedicated section (5.2) on dynamic mental-health benchmarks with…
An update on our mental health work
A Google blog post announcing changes to Gemini's handling of mental-health-related conversations, including a redesigned 'Help is available' module developed with clinical experts and a new 'one-tou…
Findings from transparency notices on AI companion apps: October 2025 (non-periodic)
Australia's eSafety Commissioner reports findings from Basic Online Safety Expectations transparency notices issued on 16 October 2025 to four AI companion providers — Chai Research Corp., Character…
System Card: Claude Opus 4.6
Anthropic's 213-page system card for Claude Opus 4.6, notable for an expanded 'user wellbeing evaluations' section covering child safety, suicide and self-harm, and eating disorders, alongside sycoph…
MiniMax Group Inc. Prospectus (Global Offering)
MiniMax Group Inc.'s prospectus for its Hong Kong Stock Exchange global offering, parent company of the AI-companion app Talkie/Xingye. The Risk Factors and Business sections disclose a staged AI saf…