Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study
A peer-reviewed development-and-validation study of ASTRA, an external system that monitors AI-mediated mental-health conversations to identify clinically relevant risk behaviors at the conversation level rather than turn by turn. It is validated against clinician-authored synthetic therapeutic dialogues across a defined set of risk categories and benchmarked against expert human raters.
Key Findings
- ASTRA was evaluated on 100 synthetic therapeutic conversations written by licensed clinicians, spanning subtle and overt risk behaviors across 8 predefined categories
- Conversation-level detection accuracy exceeded 0.90 for all risk-behavior categories
- Agreement with human expert raters was strong (Cohen's kappa 0.65-1.00)
- Results support the feasibility of independent, whole-conversation safety-monitoring systems as a complement to AI mental-health tools
Methodology Notes
Development-and-validation study; test corpus of 100 clinician-authored synthetic conversations (not real patient data) with embedded risk behaviors across 8 categories; whole-conversation rather than single-turn evaluation; benchmarked against expert human raters. Accepted 20 April 2026, published 19 May 2026 (JMIR Mental Health, vol 13, e91367).
Sources
JMIR Mental Health article (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Daniel Szoke, Ilana Hutzler, Jerry Liu, Samantha Addante, Zuhaib Akhtar, Dale L. Smith, Kirsten Dickins, Charles Small, Sarah Pridgen, Philip Held
Tags
Cite This
APA
Daniel Szoke et al. (2026). Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study. JMIR Mental Health. https://mental.jmir.org/2026/1/e91367
Related Insights
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
arXiv (Emory University-led) · 27 Oct 2025
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
arXiv (Princeton University-led) · 13 Apr 2025
Beyond Engagement: A Multidimensional Framework to Evaluate the Safe Development of Agentic AI in Mental Health
Lecture Notes in Computer Science (Springer Nature) — AI for Clinical Applications · 22 Sept 2025