Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

GPT-Live System Card

System card for GPT-Live-1 and GPT-Live-1 mini, OpenAI's full-duplex voice models that became the default voice models for paid and free ChatGPT users respectively. The card introduces voice-native safety evaluations built from real opted-in user audio and from synthetic spoken prompts, comparing the new models with the Advanced Voice Mode models they replace across sexual content, illicit behaviour, mental health, personal data, emotional reliance and self-harm. It also describes voice-specific system safeguards that can interrupt a response, play a spoken safety message, show support resources in text, or end the conversation. A change log dated 2026-08-04 records that the originally published results were generated with a mismatched backend configuration and were re-run, with the original and corrected values shown side by side.

Publisher

OpenAI

Published

8 Jul 2026

Added

today

DOI

Key Findings

  • Voice-native production prompts (adversarially selected, not prevalence-weighted), GPT-Live-1 vs Advanced Voice Mode: emotional reliance 0.82 vs 0.88 (a regression the card calls slight), mental health 0.90 vs 0.90, self-harm 0.97 vs 0.89, illicit behaviour 0.95 vs 0.74
  • GPT-Live-1 mini vs AVM mini on the same set: emotional reliance 0.78 vs 0.78, mental health 0.85 vs 0.78, self-harm 0.95 vs 0.81, sexual content 0.92 vs 0.97
  • Synthetic spoken prompts targeting policy edge cases, GPT-Live-1 vs AVM: emotional reliance 0.93 vs 0.72, mental health 0.85 vs 0.57, self-harm 0.96 vs 0.72
  • Voice-specific system safeguards check inputs and outputs as the conversation unfolds and can steer or interrupt the response, play a spoken safety message, provide support resources in text, or end the voice conversation in higher-risk cases
  • Red teaming covered child-coded voices, impersonation, self-harm, emotional reliance, scams and manipulation; sexual content, emotional reliance and self-harm were prioritised for mitigation
  • The 2026-08-04 correction reports statistically significant regressions relative to the first-published values on illicit behaviour, synthetic sexual content and gore, with OpenAI stating that none fell below its launch standards

Methodology Notes

Published 2026-07-08 at deploymentsafety.openai.com with a change log entry dated 2026-08-04; the PDF of record is 44 KB with two evaluation tables and an appendix of original versus corrected values. Production prompts are transcribed from real audio shared by opted-in users after PII scrubbing; synthetic prompts are generated from safety policies and converted to speech. Both sets are built around cases where the predecessor models were not giving ideal responses and are explicitly not a guide to production prevalence. The card does not define what the 'mental health' category covers, and the models can delegate work to text models whose safety training then applies. Card read in full from the PDF text.

Tags

system-cardopenaivoicefull-duplexemotional-reliancecorrectionadvanced-voice-mode

Cite This

APA

OpenAI. (2026). GPT-Live System Card. https://deploymentsafety.openai.com/gpt-live