GPT-6 Sol and GPT-6 Luna: October 2026 update
System card for the October 2026 release of GPT-6 Sol and GPT-6 Luna, which replace GPT-5.6 Sol and GPT-5.6 Luna across all free and paid ChatGPT plans globally (Codex and ChatGPT Work keep the September versions). It reports disallowed-content and under-18 safe-completion evaluations, jailbreak and prompt-injection robustness, HealthBench and MentalHealthBench scores, dynamic multi-turn mental-health, emotional-reliance and self-harm evaluations, hallucination and alignment evaluations, and Preparedness Framework assessments. Safety evaluations are run at the lowest reasoning setting and without system-level safeguards.
Publisher
OpenAI
Published
7 Oct 2026
Added
today
DOI
—
Key Findings
- Production safe-completion benchmarks (Table 1): GPT-6 Sol (October) shows a statistically significant regression on standard self-harm (0.896 against 0.934 for GPT-5.6 Sol (August)); GPT-6 Luna (October) regresses on self-harm (0.901 against 0.932), gore (0.812 against 0.867) and sexual content (0.899 against 0.971). Sexual content involving minors fell to 0.887 (Sol) and 0.829 (Luna) from 0.922 and 0.966. OpenAI's manual review calls the violations borderline but generally safe, says the models are more willing to answer informational self-harm questions while directing users to professional resources, and notes that the evaluations exclude system-level interventions such as Trusted Contact, localised crisis helplines and Parental Controls.
- Under-18 evaluations (Table 2): emotional reliance fell from 0.921 to 0.770 (Sol) and from 0.927 to 0.734 (Luna); sexual content from 0.971 to 0.933 and from 0.984 to 0.904; age-restricted goods, services and dangerous challenges from 0.865 to 0.799 and from 0.857 to 0.814; eating disorders from 0.808 to 0.772 and from 0.810 to 0.776; under-18 self-harm from 0.987 to 0.977 and from 0.977 to 0.974. OpenAI reports the regressions on age-restricted content, sexual content and emotional reliance as statistically significant, attributes the emotional-reliance drop to an evaluation that is over-sensitive to benign nicknames such as 'bro' or 'bestie' when the user requests them, and says an additional classifier-based block on self-harm, sexual content and gore for teens is not captured in these results.
- MentalHealthBench (Table 6; 1,215 synthetic conversations, 0 to 100): overall 48.81 (Sol) and 51.73 (Luna) against 47.53 and 45.11 for the GPT-5.6 (August) pair and 53.93 for GPT-5.5 Instant (June); high-acuity 53.70 and 56.15; emergent 47.92 and 48.98 against 44.30 and 44.12.
- Dynamic multi-turn benchmarks with adversarial user simulations (Table 7; share of assistant messages that do not violate policy): mental health 0.996 (Sol) and 0.991 (Luna); emotional reliance 0.961 for both; self-harm 0.889 and 0.884 against 0.901 and 0.911 for the GPT-5.6 (August) pair; OpenAI states that none of these differences is statistically significant.
- HealthBench (Table 5; length-adjusted, with unadjusted scores in parentheses): HealthBench Professional 55.2 (62.2) for Sol and 48.2 (57.8) for Luna, up from 54.0 and 44.1; HealthBench 52.1 (55.7) and 50.4 (58.5) against 55.0 and 53.3. The card states that HealthBench is approaching a noise ceiling for frontier models and recommends HealthBench Professional.
- Instruction-hierarchy robustness 99.99% (Sol) and 99.79% (Luna); indirect prompt-injection robustness 97.13% and 95.80%; multi-turn jailbreak robustness slightly below the September versions with overlapping confidence intervals. Under the Preparedness Framework the release is treated as High capability in Cybersecurity and in Biological and Chemical domains, below the High threshold for AI Self-Improvement, with the same safeguards as GPT-5.6.
Methodology Notes
Developer self-report. The disallowed-content, under-18 and dynamic evaluation sets are built around difficult production-derived and long-tail cases; the card states that error rates are not representative of average traffic, that safety evaluations are measured at the lowest reasoning deployment setting without system-level safeguards, and that comparison values for earlier models come from their latest versions. Significance statements are OpenAI's. Tables 1, 2, 3 and 7 are client-rendered charts whose values were read from the chart data embedded in the page (six model columns: GPT-5.5 Instant May and June, GPT-5.6 Luna and Sol August, GPT-6 Sol and Luna October); Tables 5 and 6 are HTML tables. MentalHealthBench scores average four responses per task over 1,215 tasks (650 non-acute, 221 high-acuity, 344 emergent). The versions released on 2026-10-07 are labelled October; the September versions remain in Codex and ChatGPT Work. Date from the page header and the Deployment Safety Hub RSS item. A 24-page PDF is linked from the page.
Sources
OpenAI Deployment Safety Hub system card(opens in a new tab) (primary)
System card PDF (24 pp)(opens in a new tab) (7 Oct 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Tags
Cite This
APA
OpenAI. (2026). GPT-6 Sol and GPT-6 Luna: October 2026 update. https://deploymentsafety.openai.com/gpt-6-october/
Related Insights
GPT-6 Astra System Card
OpenAI · 3 Sept 2026
GPT-5.6 System Card
OpenAI · 9 Jul 2026
GPT-5.6 – August Updates
OpenAI · 6 Aug 2026
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
OpenAI · 23 Sept 2026
HealthBench: Evaluating Large Language Models Towards Improved Human Health
OpenAI (arXiv preprint) · 13 May 2025
ChatGPT for Teens: AI Risk Assessment
Common Sense Media (Youth AI Safety Institute) · 7 Oct 2026
System Card: Claude Haiku 5.5
Anthropic · 7 Oct 2026