Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

31 artifacts matching

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

18 Sept 2026 OpenAI Lab publication

Lab publication

An Australian Youth Safety Blueprint

Seven-page policy document in which OpenAI sets out six pillars it proposes for Australian youth-AI policy: recognising positive uses and AI literacy in education; privacy-preserving age assurance; u…

8 Sept 2026 OpenAI (OpenAI Group PBC) Lab publication

Lab publication

OpenAI's AI and Teen Development Research Grant Program

A funding call from OpenAI for independent research on how generative AI use relates to adolescent development, including safeguards and design choices. The programme opened 8 September 2026 with a d…

3 Sept 2026 OpenAI Lab publication

Lab publication

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…

31 Aug 2026 Transluce Benchmark / dataset

Benchmark / dataset

Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)

Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…

25 Aug 2026 arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) Benchmark / dataset

Benchmark / dataset

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…

24 Aug 2026 arXiv (preprint) Preprint

Preprint

AI emotional support is better only when chosen, but shifts preferences even when it is not

Three experiments plus a 28-day field study examining how people choose between human and AI emotional support and what happens when the support they receive does not match what they chose. Participa…

18 Aug 2026 OpenAI Lab publication

Lab publication

Introducing ChatGPT for Teens: Built for learning, backed by protections

OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…

18 Aug 2026 OpenAI Lab publication

Lab publication

OpenAI Model Spec (August 18, 2026 version)

The current version of OpenAI's published behavioural specification for its models. Relative to the October 2025 version, the December 2025 release added Under-18 Principles for users aged 13 to 17 a…

6 Aug 2026 OpenAI Lab publication

Lab publication

GPT-5.6 – August Updates

System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…

9 Jul 2026 OpenAI Lab publication

Lab publication

GPT-5.6 System Card

General-availability system card for the GPT-5.6 family (Sol, the flagship; Terra, a lower-cost model; Luna, the fastest), published alongside the models' broad rollout. Retains Section 5.2's dynamic…

8 Jul 2026 OpenAI Lab publication

Lab publication

GPT-Live System Card

System card for GPT-Live-1 and GPT-Live-1 mini, OpenAI's full-duplex voice models that became the default voice models for paid and free ChatGPT users respectively. The card introduces voice-native s…

26 Jun 2026 OpenAI Lab publication superseded

Lab publication

GPT-5.6 Preview System Card

OpenAI's system card for the GPT-5.6 preview, a family of three models (Sol, the flagship; Terra, a lower-cost option; and Luna, the fastest/most cost-efficient), released in a limited preview ahead…

18 Jun 2026 OpenAI (Alignment Research Blog; arXiv preprint) Lab publication

Lab publication

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

OpenAI alignment research asking whether reinforcement learning on realistic conversations that reward beneficial traits (truthfulness, fairness, risk awareness, corrigibility, epistemic humility, co…

14 May 2026 OpenAI Lab publication

Lab publication

Helping ChatGPT better recognize context in sensitive conversations

OpenAI post describing safety updates that let ChatGPT recognize risk emerging over the course of a conversation rather than judging each message alone, focused on suicide, self-harm and harm-to-othe…

1 May 2026 Harvard Business School Preprint

Preprint

Benchmarking the Safety of General Purpose Large Language Models for Suicide Risk Detection and Response

Harvard Business School working paper applying the open-source VERA-MH LLM-as-judge evaluation framework to ten general-purpose models from OpenAI, Anthropic, Google DeepMind, and xAI, scoring suicid…

23 Apr 2026 OpenAI Lab publication

Lab publication

GPT-5.5 System Card

OpenAI's system card for GPT-5.5, published on its Deployment Safety Hub, documenting safety evaluations for the model. It includes a dedicated section (5.2) on dynamic mental-health benchmarks with…

1 Mar 2026 Association for Computational Linguistics (Proceedings of the 1st Workshop on Linguistic Analysis for Health, HeaLing 2026) Peer-reviewed

Peer-reviewed

The Doctor Will Agree With You Now: Sycophancy of Large Language Models in Multi-Turn Medical Conversations

Evaluates sycophancy in ten language models from OpenAI, Google and Anthropic under a four-turn escalatory pushback protocol on open-ended diagnostic cases (MedCaseReasoning) and clear-answer biomedi…

23 Feb 2026 Nature Medicine (Springer Nature); Icahn School of Medicine at Mount Sinai Peer-reviewed

Peer-reviewed

ChatGPT Health performance in a structured test of triage recommendations

Brief Communication reporting a structured stress test of ChatGPT Health, the consumer health feature OpenAI launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains…

1 Nov 2025 OpenAI Framework

Framework

Protecting Teen ChatGPT Users: OpenAI's Teen Safety Blueprint

OpenAI's public commitments framework for protecting teenage ChatGPT users, covering age prediction, age-appropriate response policies, and parental controls, developed with input from policymakers (…