Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The participant-facing sections report production safety benchmarks, a set of six under-18 evaluations, HealthBench scores, and dynamic multi-turn mental-health, emotional-reliance and self-harm evaluations with adversarial user simulations, each compared against GPT-5.6 Sol and two earlier models. The card states that under-18 safety was integrated into post-training for the first time and describes 'safety context', a system-level safeguard that combines safety signals across conversations for ambiguous-intent self-harm and violence scenarios.

Publisher

OpenAI

Published

3 Sept 2026

Added

today

DOI

Key Findings

  • Under-18 evaluations (scored with U18 protections in place; higher is better), GPT-6 Astra vs GPT-5.6 Sol: eating disorders 0.921 vs 0.710, age-restricted goods and dangerous challenges 0.918 vs 0.719, sexual content 0.991 vs 0.929, emotional reliance 0.944 vs 0.931, self-harm 0.995 vs 0.982, gore 0.898 vs 0.785
  • Dynamic multi-turn benchmarks with adversarial user simulations (share of assistant messages not violating policy): mental health 1.000 (Sol 0.991; GPT-5.5-thinking 0.820), emotional reliance 0.993 (0.953; 0.915), self-harm 0.989 (0.856; 0.868)
  • Production benchmarks with challenging prompts: self-harm (standard) 0.992 vs 0.945 for Sol, sexual/minors 0.974 vs 0.973, violent illicit behaviour 0.990 vs 0.934; the card also claims fewer refusals of harmless requests and fewer unnecessarily judgmental caveats
  • For users known or predicted to be under 18 the model is trained not to engage in romantic roleplay, encourage age-restricted challenges, or position itself as a substitute for real-world relationships; across all ages the model 'should not encourage exclusive or dependency-forming relationships'
  • 'Safety context' is described as a system-level safeguard that brings together safety signals across conversations, evaluated on self-harm and violence scenarios where no single message reveals intent; gains are shown only as a figure, and for users flagged as potentially high risk the model can shift its refusal boundary to be more conservative
  • Length-adjusted HealthBench Professional 63.4 (+2.9 vs Sol), HealthBench 58.1 (+1.1), HealthBench Hard 36.3 (+3.2), HealthBench Consensus 95.8 (+0.3); OpenAI states HealthBench is approaching a noise ceiling and recommends HealthBench Professional

Methodology Notes

Published 2026-09-03 at deploymentsafety.openai.com; the HTML page strips table bodies, so all figures were read from the PDF of record (10.2 MB, about 60 pages plus appendices). Evaluations are internal OpenAI benchmarks run without system-level safeguards except where stated; production and U18 sets are described as deliberately difficult, built around cases where prior models were not yet giving ideal responses, and 'not representative of average production traffic'. Comparison values for earlier models are re-run on their current versions. HealthBench scores are length-adjusted with stated penalties per 500 characters beyond 2,000. External evaluators named include UK AISI, Apollo Research, SecureBio and Irregular, for cyber and alignment work rather than the mental-health sections. No dataset, prompt or grader release. A same-day 'Safety overview: GPT-6 Astra' post duplicates section 1 of the card.

Tags

system-cardgpt-6openaiteen-safetyu18emotional-reliancehealthbenchsafety-contextmulti-turn

Cite This

APA

OpenAI. (2026). GPT-6 Astra System Card. https://deploymentsafety.openai.com/gpt-6-astra