Muse Spark Safety & Preparedness Report
Meta's safety and preparedness report for Muse Spark, the model behind the Meta AI assistant. It presents catastrophic-risk evaluations under Meta's Advanced AI Scaling Framework, jailbreak and agent-misuse tests, and a model-behaviour section that includes an internal sycophancy benchmark, honesty, harm aversion and contextual privacy, followed by a content-safety section describing violation-rate evaluations, a separate youth safety taxonomy for users under 18 and the suicide and self-harm policy. A Muse Spark 1.1 Evaluation Report (July 2026) updates the behaviour and jailbreak results for the model version then powering Meta AI.
Publisher
Meta (arXiv preprint)
Published
14 May 2026
Added
today
DOI
—
Key Findings
- Internal sycophancy benchmark (adversarially curated from real transcripts where a model was sycophantic and from simulated users; includes endorsing harmful framings in emotionally charged situations and failing to challenge dangerous beliefs): Meta AI in Thinking mode 50.1%, against Claude Opus 4.6 at 50.9% and GPT-5.4 at 45.4%; the Muse Spark model without system-level mitigations 62.9%, second highest after Gemini 3.1 Pro (65.6%), falling to 57.7% with additional reasoning. The authors state the benchmark does not measure how often sycophancy occurs in average conversations.
- Muse Spark 1.1 Evaluation Report (2026-07-09): sycophancy rate 49.2% for Muse Spark 1.1 against 57.9% for Muse Spark 1.0, 45.5% for GPT-5.5, 32.4% for Claude Opus 4.8 and 65.6% for Gemini 3.1 Pro; excessive anti-sycophancy (inappropriate pushback) rose from 10.7% to 13.9%, against 23.4% for Claude Opus 4.8.
- Content safety is measured as violation rates judged by LLM-based safety judges calibrated against human annotation; a separate youth safety taxonomy governs behaviour for users under 18 and is evaluated before launch, benchmarked against other assistants that offer a dedicated youth experience. The report states that Meta AI met internal thresholds in every category but publishes no quantitative content-safety results.
- Suicide and self-harm content is placed in the highest severity tier because Meta AI is integrated into messaging and social platforms where it may encounter users in distress; policy requires refusing content that promotes or instructs self-harm and directing users to crisis support resources when intent may be indicated.
- Pre-mitigation Chemical and Biological capability was assessed as likely reaching the Framework's 'high risk' category; Meta released Muse Spark as the underlying model of Meta AI after mitigations.
Methodology Notes
Developer self-report. Behaviour benchmarks are internal and adversarially constructed, so rates indicate relative resistance under adversarial testing, not prevalence; comparisons with other vendors used their APIs with system mitigations where stated. Content-safety thresholds and results are not published. The 1.0 and 1.1 reports give different Muse Spark 1.0 sycophancy figures (62.9% without system mitigations and 50.1% for Meta AI Thinking in the 1.0 report; 57.9% in the 1.1 table), so configurations differ between reports. Dates: arXiv v1 submitted 2026-05-14; Muse Spark launched in Meta AI on 2026-04-08 (Meta newsroom post, which does not link the report), so the Meta-hosted PDF may predate the arXiv version but its first publication date was not established; the 1.1 report is dated 2026-07-09 on its title page. Author list abridged (120 authors in the arXiv citation metadata).
Sources
arXiv preprint(opens in a new tab) (primary)
Meta-hosted PDF (direct HTTP 400; Wayback capture 2026-09-15)(opens in a new tab) (14 May 2026)
Muse Spark 1.1 Evaluation Report (direct HTTP 400; Wayback capture 2026-10-04)(opens in a new tab) (9 Jul 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan Wang, Miles Turpin, Felix Binder, Nathaniel Li, Kate Plawiak, Eric Michael Smith
Tags
Cite This
APA
Cristina Menghini et al. (2026). Muse Spark Safety & Preparedness Report. Meta (arXiv preprint). https://arxiv.org/abs/2606.12429
Related Insights
Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI
Meta · 16 Jul 2026
Evaluating AI Safety in Teen Conversations
Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine · 23 Sept 2026
GPT-6 Astra System Card
OpenAI · 3 Sept 2026
Claude Opus 5.5 System Card
Anthropic · 22 Sept 2026