31 artifacts matching
Benchmark / dataset
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…
Lab publication
An Australian Youth Safety Blueprint
Seven-page policy document in which OpenAI sets out six pillars it proposes for Australian youth-AI policy: recognising positive uses and AI literacy in education; privacy-preserving age assurance; u…
Lab publication
OpenAI's AI and Teen Development Research Grant Program
A funding call from OpenAI for independent research on how generative AI use relates to adolescent development, including safeguards and design choices. The programme opened 8 September 2026 with a d…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Benchmark / dataset
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…
Benchmark / dataset
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…
Preprint
AI emotional support is better only when chosen, but shifts preferences even when it is not
Three experiments plus a 28-day field study examining how people choose between human and AI emotional support and what happens when the support they receive does not match what they chose. Participa…
Lab publication
Introducing ChatGPT for Teens: Built for learning, backed by protections
OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…
Lab publication
OpenAI Model Spec (August 18, 2026 version)
The current version of OpenAI's published behavioural specification for its models. Relative to the October 2025 version, the December 2025 release added Under-18 Principles for users aged 13 to 17 a…
Lab publication
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Lab publication
GPT-5.6 System Card
General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…
Lab publication
GPT-Live System Card
System card for GPT-Live-1 and GPT-Live-1 mini, OpenAI's full-duplex voice models that became the default voice models for paid and free ChatGPT users respectively. The card introduces voice-native s…
Lab publication
GPT-5.6 Preview System Card
OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…
Lab publication
Reinforcement Learning Towards Broadly and Persistently Beneficial Models
OpenAI alignment research asking whether reinforcement learning on realistic conversations that reward beneficial traits (truthfulness, fairness, risk awareness, corrigibility, epistemic humility, co…
Lab publication
Helping ChatGPT better recognize context in sensitive conversations
OpenAI post describing safety updates that let ChatGPT recognize risk emerging over the course of a conversation rather than judging each message alone, focused on suicide, self-harm and harm-to-othe…
Preprint
Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response
Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…
Lab publication
GPT-5.5 System Card
OpenAI's system card for GPT-5.5, published on its Deployment Safety Hub, documenting safety evaluations for the model. It includes a dedicated section (5.2) on dynamic mental-health benchmarks with…
Peer-reviewed
The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations
Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedi…
Peer-reviewed
ChatGPT Health performance in a structured test of triage recommendations
Brief Communication reporting a structured stress test of ChatGPT Health, the consumer health feature OpenAI launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains…
Framework
Protecting Teen ChatGPT Users: OpenAI's Teen Safety Blueprint
OpenAI's public commitments framework for protecting teenage ChatGPT users, covering age prediction, age-appropriate response policies, and parental controls, developed with input from policymakers (…