What model-makers publish
From the Labs
System cards, usage studies, and safety research published by the AI labs themselves — primary evidence of how frontier models behave and how their makers measure it.
33 entries, newest first
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Shieldstral
Technical report introducing Shieldstral, a 3B-parameter open-weights (Apache 2.0) policy-adaptive multimodal safety classifier from Mistral AI. Content moderation is reformulated as binary question-…
System Card: Claude Opus 5
System card for Claude Opus 5, an upgrade to Claude Opus 4.8, with a dedicated mental-health evaluation section (4.3) covering single- and multi-turn suicide/self-harm handling and disordered eating,…
Mila's Suicide Prevention Guardrail: Lightweight, Open Source Safeguards for Real-Time Detection
Joint beta release by Mila's AI Safety Studio and ROOST of an open-weights (Apache-2.0) output-moderation classifier for suicide and self-harm content in chatbot responses. The model is a fine-tuned…
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
System Card: Claude Sonnet 5
Anthropic's system card for Claude Sonnet 5, an upgrade to Sonnet 4.6. Reports that hallucination and sycophancy are qualitatively 'markedly improved' relative to Sonnet 4.6, while 'wet blanket' resp…
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
System Card: Claude Fable 5 & Claude Mythos 5
Anthropic's 317-page system card for Claude Fable 5 (general-release) and Claude Mythos 5 (restricted trusted-access release), the two safeguard configurations of a new frontier model. Alongside exte…
System Card: Claude Opus 4.8
Anthropic's 246-page system card for Claude Opus 4.8, including single-turn and multi-turn suicide/self-harm evaluations and qualitative expert review of the model's crisis-conversation handling rela…
How people ask Claude for personal guidance
An Anthropic research analysis of roughly 38,000 personal-guidance conversations (sampled from about 1M) covering significant life decisions across health/wellness, career, relationships, and persona…
GPT-5.5 System Card
OpenAI's system card for GPT-5.5, published on its Deployment Safety Hub, documenting safety evaluations for the model. It includes a dedicated section (5.2) on dynamic mental-health benchmarks with…
Emotion Concepts and their Function in a Large Language Model
Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…
An update on our mental health work
A Google blog post announcing changes to Gemini's handling of mental-health-related conversations, including a redesigned 'Help is available' module developed with clinical experts and a new 'one-tou…
System Card: Claude Opus 4.6
Anthropic's 213-page system card for Claude Opus 4.6, notable for an expanded 'user wellbeing evaluations' section covering child safety, suicide and self-harm, and eating disorders, alongside sycoph…
MiniMax Group Inc. Prospectus (Global Offering)
MiniMax Group Inc.'s prospectus for its Hong Kong Stock Exchange global offering, parent company of the AI-companion app Talkie/Xingye. The Risk Factors and Business sections disclose a staged AI saf…