223 artifacts
Introducing ChatGPT for Teens: Built for learning, backed by protections
OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…
The new first listener: Daydreaming styles and self-compassion predict adolescent disclosure to AI, differently in ADHD
Cross-sectional survey of 2,115 adolescents and young people in the United States and Hong Kong examining how daydreaming styles and self-compassion relate to disclosing inner thoughts to an AI chatb…
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…
AI Chatbots as Companions: Overview, Uses, and Considerations for Congress
A Congressional Research Service report for the 119th Congress on AI chatbots used for companionship. It separates purpose-built companion chatbots from general-purpose chatbots used for companionshi…
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
Longitudinal qualitative evaluation of whether mainstream chatbots exacerbate an unfolding psychotic process. Fifteen widely used models were prompted across 30 days with the same 30-message script s…
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
Preregistered experiment in which 1,500 UK adults each held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical across conditions and only the disc…
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of…
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
Pre-registered four-week longitudinal study (N=72, 182,451 lines of conversation) in which participants conversed with ChatGPT-4o either under a relational system prompt or unmodified, analysed throu…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Interview study with 19 practitioners who work directly with youth in vulnerable situations — social workers, therapists and psychologists — asking them to assess chatbot responses to risky situation…
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Peer-reviewed Nature Medicine study introducing SIM-VAIL (simulated vulnerability-amplifying interaction loops), a clinically validated framework for auditing chatbot behavior in mental-health contex…
Measuring and Detecting Harmful AI Sycophancy
Large-scale measurement and detection study of preference-induced stance reversal (PSRS) — the harmful form of sycophancy where a model abandons a correct or safe position after user pushback. Builds…
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Evaluation protocol testing chatbots' tendencies to exhibit behaviors linked to promoting user delusions, grounded in real conversation histories rather than synthetic scenarios. Models are prompted…
Interaction with AI companions and psychological well-being
Stanford-led study of 1,131 adult Character.AI users combining survey self-report with donated chat transcripts from 244 of them, analyzed with LLM-assisted methods against the Comprehensive Inventor…
Talking to machines: Children's experiences with AI assistants and companions
National survey report from Australia's eSafety Commissioner on children's use of AI assistants and companions, based on 1,950 children aged 10-17 surveyed in February-March 2026. Covers prevalence a…
Adolescent AI Use and Evaluations: Prospective Bidirectional Associations with Internalizing Symptoms
Two-wave longitudinal cohort study of 2,342 students in grades 6-8 across 22 middle schools in eight districts of a large southeastern US state, measuring AI chatbot use, use frequency and evaluation…
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
Large-scale factorial study of medical sycophancy — models abandoning correct medical answers under user pushback — crossing four conversational factors with five open-weight models over 500 MedQuAD-…
Real-world use of large language models for mental health in 2024
Survey of 1,871 US adults conducted between August and October 2024, using stratified sampling across age, sex and race/ethnicity to approximate national demographics, measuring how many people use g…