Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

Model Card: Grok 4.7

Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's safety evaluations for jailbreak robustness, general refusals in six languages, child safety, self-harm and crisis handling, honesty under pressure and sycophancy, each compared with Grok 4.5 and Grok 4.6, and describes a layered safeguard stack including runtime filters for self-harm and child sexual abuse material on some surfaces.

Publisher

xAI (styled SpaceXAI in the card)

Published

21 Sept 2026

Added

today

DOI

Key Findings

  • Self-harm suite compliance (failures, where the suite also counts refusing without redirecting the user to help or missing implied crisis intent) rose across releases: 0.50% for Grok 4.5, 0.84% for Grok 4.6 and 1.05% for Grok 4.7; the card states that self-harm compliance is higher than for 4.6
  • Child-safety multi-turn evaluation (child sexual exploitation, including CSAM) reports 0.0% compliance for Grok 4.5, 4.6 and 4.7
  • General refusal compliance across English, Spanish, Chinese, Japanese, Arabic and Russian prompts: 1.10% (4.5), 0.93% (4.6), 1.10% (4.7)
  • Jailbreak compliance: standard battery 0.01% (from 0.04% for 4.6), StrongREJECT 2.0% (from 3.9%), long-horizon multi-turn attacks including Crescendo 0.65% (from 1.0%)
  • Sycophancy (abandoning a correct factual answer under a confident wrong user claim) 0.03% (4.6: 0.04%); MASK-Rectified dishonesty 0.00% (4.6: 1.90%)
  • Grok 4.7 uses a new, larger base model with a June 2026 pretraining cutoff and supplemental data to August 2026; HealthBench Professional score 56.7%

Methodology Notes

Vendor self-report. All safety evaluations are internal suites graded by a model-based grader; the jailbreak section cites StrongREJECT and Crescendo. Compliance figures are shares of should-refuse items where disallowed assistance was given (lower is better). No confidence intervals, prompt counts or evaluation dates are given for the safety suites. Date from the card cover ('September 21, 2026, Revision: 2026-09-21') and PDF metadata (created 2026-09-21 15:25 UTC). The PDF is served from media.x.ai under a hashed filename and was fetched directly (HTTP 200, 332,355 bytes, 30 pages) and read; x.ai returns 403 to automated fetches, so the announcement page was read from a Wayback raw capture.

Tags

model-cardgrokxaispacexaiself-harm-refusalssycophancyjailbreakschild-safety

Cite This

APA

xAI (styled SpaceXAI in the card). (2026). Model Card: Grok 4.7. https://media.x.ai/v1/website/4p7card-5eccc980.pdf