Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability benchmarks, but carries dedicated safety sections: general output safety, child safety, CBRN refusals, jailbreak robustness, a mental-health section covering self-harm and crisis handling, and behavior propensities (honesty under pressure, sycophancy). The card self-reports regressions versus Grok 4.5 on self-harm refusal compliance, dishonesty under pressure, and sycophancy.
Publisher
xAI
Published
12 Aug 2026
Added
1 month ago
DOI
—
Key Findings
- Self-harm suite compliance regressed from 0.50% (Grok 4.5) to 0.84% (Grok 4.6, lower is better) in the revised card of 2026-08-17; the original 2026-08-12 card reported 3.7% for Grok 4.6 before xAI's changelog entry 'Corrected eval results on HackerBench v0.2, Self-harm, MASK, LAB'. The suite covers self-harm and crisis prompts including longer-horizon conversations and counts a refusal without redirection to help as a failure
- MASK-Rectified dishonesty regressed from 0.67% to 1.90% (revised card; the original card reported 3.8%), and sycophancy (abandoning a correct answer under misleading user-supplied context) from 0.01% to 0.04%
- Child-safety/CSAM compliance unchanged at 0.0% and general refusals improved from 1.1% to 0.93%; jailbreak robustness reported at 0.04% (standard), 3.9% (StrongReject), 1.0% (long-horizon)
- Reported results cover the API, Cursor, and Grok Build channels; consumer surfaces (Grok in X, standalone apps) are stated to receive the model at a later date
Methodology Notes
Published 2026-08-12 (card cover date and revision stamp), 36-page PDF at media.x.ai. The card self-identifies its publisher as "SpaceXAI", stated in a footnote to be a doing-business-as name of XAI LLC used interchangeably with xAI. Safety results come from internal evaluation suites using compliance-style metrics (lower is better) applied to Grok 4.5 and 4.6 side by side; internal evaluations are not itemized in the references. Verified by fetching and reading the primary PDF directly. REVISION (recorded 2026-09-03): the PDF at the same media.x.ai URL now carries 'Revision: 2026-08-17' with a changelog entry 'Corrected eval results on HackerBench v0.2, Self-harm, MASK, LAB'; the self-harm compliance figure for Grok 4.6 changed from 3.7% to 0.84% and MASK-Rectified dishonesty from 3.8% to 1.90% (Grok 4.5 baselines unchanged at 0.50% and 0.67%). The document was revised in place with no version-specific URL; the 2026-08-13 Wayback snapshot preserves the original figures. Revised PDF fetched and read 2026-09-03.
Sources
Grok 4.6 model card PDF (media.x.ai)(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Tags
Cite This
APA
xAI. (2026). Model Card: Grok 4.6. https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf
Related Insights
GPT-5.6 System Card
OpenAI · 9 Jul 2026
System Card: Claude Opus 5
Anthropic · 24 Jul 2026
PIPEDA Findings #2026-004: Commissioner-Initiated Complaint Concerning X Corp. and X.AI LLC
Office of the Privacy Commissioner of Canada (OPC) · 11 Jun 2026
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic · 1 Sept 2026