Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

105 artifacts matching

8 Sept 2026 arXiv (Grow Therapy; Stanford University School of Medicine) Preprint

Preprint

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…

8 Sept 2026 World Journal of Psychiatry (Baishideng Publishing Group) Peer-reviewed

Peer-reviewed

Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support

Prospective, non-randomized, waitlist-controlled pilot at a junior high school in Zhejiang Province, China, of 'Duoduo', an acceptance-and-commitment-therapy-informed multimodal generative-AI chatbot…

5 Sept 2026 arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 Preprint

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…

3 Sept 2026 Character.AI (Character Technologies) Lab publication

Lab publication

Continuing To Build Upon Our Safety Priorities

A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…

3 Sept 2026 OpenAI Lab publication

Lab publication

GPT-6 Astra System Card

System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

1 Sept 2026 Frontiers in Psychology (Frontiers Media) Peer-reviewed

Peer-reviewed

Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions

2 by 2 randomized vignette experiment with 768 Chinese college students testing whether a privacy-assurance message and a professional-boundary warning in a mental-health chatbot interface change per…

1 Sept 2026 Microsoft (Office of Responsible AI) Lab publication

Lab publication

2026 Responsible AI Transparency Report

Microsoft's third annual responsible-AI transparency report, covering July 2025 to June 2026. One of its four 2026 trends is people turning to conversational AI for personal advice, health questions…

29 Aug 2026 AI & SOCIETY (Springer) Peer-reviewed

Peer-reviewed

Benevolent Gravity: the lethal structure inherent in conversational AI design principles

Analytical paper arguing that the two dominant explanations for fatalities linked to conversational AI — safety-filter failure and commodified intimacy — are structurally insufficient, because a subs…

29 Aug 2026 ACM (Proceedings of Mensch und Computer 2026) Peer-reviewed

Peer-reviewed

Beyond Manipulation: How Users Perceive Harmful AI Chatbot Interactions

Mixed-methods study (N = 100) in which participants recalled a positive, an inappropriate or a manipulative chatbot interaction. Exploratory factor analysis of an adapted perceived-manipulation quest…

28 Aug 2026 Anthropic Lab publication

Lab publication

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…

26 Aug 2026 arXiv (Nanyang Technological University; National University of Singapore) Benchmark / dataset

Benchmark / dataset

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…

26 Aug 2026 arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo) Benchmark / dataset

Benchmark / dataset

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark ho…

18 Aug 2026 U.S. Food and Drug Administration, Center for Devices and Radiological Health Regulator study

Regulator study

Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback

A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…

18 Aug 2026 OpenAI Lab publication

Lab publication

Introducing ChatGPT for Teens: Built for learning, backed by protections

OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…

17 Aug 2026 Journal of Psychopathology and Clinical Science (American Psychological Association) Peer-reviewed

Peer-reviewed

A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)

The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…

12 Aug 2026 xAI Lab publication

Lab publication

Model Card: Grok 4.6

36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…

6 Aug 2026 OpenAI Lab publication

Lab publication

GPT-5.6 – August Updates

System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…

6 Aug 2026 PsyArXiv (Universidad Francisco de Vitoria; Durham University) Preprint

Preprint

Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns

Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…

2 Aug 2026 arXiv (Universitat de Barcelona) Preprint

Preprint

Same violence, different answer: how AI responds to coercive control against women across languages

Cross-language audit of how conversational AI responds to a coercive-control disclosure. One scripted scenario — a woman whose partner tracks her phone asks for help writing a self-blaming letter acc…