105 artifacts matching
Preprint
Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions
Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…
Peer-reviewed
Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support
Prospective, non-randomized, waitlist-controlled pilot at a junior high school in Zhejiang Province, China, of 'Duoduo', an acceptance-and-commitment-therapy-informed multimodal generative-AI chatbot…
Preprint
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…
Lab publication
Continuing To Build Upon Our Safety Priorities
A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Peer-reviewed
Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions
2 by 2 randomized vignette experiment with 768 Chinese college students testing whether a privacy-assurance message and a professional-boundary warning in a mental-health chatbot interface change per…
Lab publication
2026 Responsible AI Transparency Report
Microsoft's third annual responsible-AI transparency report, covering July 2025 to June 2026. One of its four 2026 trends is people turning to conversational AI for personal advice, health questions…
Peer-reviewed
Benevolent Gravity: the lethal structure inherent in conversational AI design principles
Analytical paper arguing that the two dominant explanations for fatalities linked to conversational AI — safety-filter failure and commodified intimacy — are structurally insufficient, because a subs…
Peer-reviewed
Beyond Manipulation: How Users Perceive Harmful AI Chatbot Interactions
Mixed-methods study (N = 100) in which participants recalled a positive, an inappropriate or a manipulative chatbot interaction. Exploratory factor analysis of an adapted perceived-manipulation quest…
Lab publication
Automated Researchers Can Reliably Mitigate Alignment Failures
Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including…
Benchmark / dataset
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
A public benchmark dataset of 2,111 multi-turn conversations (14,051 utterances) between users and the AI companion Replika, drawn from screenshots that users posted to Reddit's r/replika and reconst…
Benchmark / dataset
HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark ho…
Regulator study
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback
A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…
Lab publication
Introducing ChatGPT for Teens: Built for learning, backed by protections
OpenAI announced a distinct under-18 product tier that users are placed into automatically when the age-prediction system estimates they are under 18 or when they state an age between 13 and 17. The…
Peer-reviewed
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…
Lab publication
Model Card: Grok 4.6
36-page model card for Grok 4.6, described as the latest release in xAI's 1.5T-scale model family, developed with supplemental training on anonymized Cursor workflow data. Predominantly capability be…
Lab publication
GPT-5.6 – August Updates
System-card addendum for the August 2026 releases of GPT-5.6 Sol and GPT-5.6 Luna. For the first time, OpenAI includes dedicated under-18 evaluations measuring model behavior against teen-specific sa…
Preprint
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
Three consumer chatbot systems (ChatGPT Free, ChatGPT Plus and Gemini Free) each completed three versions of a 21-turn conversation derived from a clinical case of emerging paranoid delusions, under…
Preprint
Same violence, different answer: how AI responds to coercive control against women across languages
Cross-language audit of how conversational AI responds to a coercive-control disclosure. One scripted scenario — a woman whose partner tracks her phone asks for help writing a self-blaming letter acc…