EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
A multi-agent framework for evaluating and mitigating mental-health harm in interactions with character chatbots. EmoEval simulates virtual users — including those portraying mentally vulnerable individuals — and scores their state with clinical instruments; EmoGuard acts as an intermediary that monitors mental status, predicts potential harm, and provides corrective feedback.
Publisher
arXiv (Princeton University-led)
Published
13 Apr 2025
Added
1 week ago
DOI
—
Key Findings
- Emotionally engaging character-chatbot dialogues produced psychological deterioration in more than 34.4% of simulated vulnerable-user runs.
- A dedicated safeguarding agent (EmoGuard) significantly reduced deterioration rates.
- Validated clinical instruments (e.g., PHQ-9, PDI, PANSS) can be used to quantify AI-induced changes in simulated user mental state.
Methodology Notes
Preprint (arXiv 2504.09689; v1 2025-04-13, revised 2025-04-29). Multi-agent simulation-and-safeguarding method; simulated-user harm quantified with clinical instruments. Title, authors, and date verified via the arXiv abstract page and arXiv API. ~15 months old and not confirmed peer-reviewed at time of logging — watch for a published version.
Sources
arXiv preprint (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Jiahao Qiu, Yinghui He, Xinzhe Juan, Yimin Wang, Yuhan Liu, Zixin Yao, Yue Wu, Xun Jiang, Ling Yang, Mengdi Wang
Tags
Cite This
APA
Jiahao Qiu et al. (2025). EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety. arXiv (Princeton University-led). https://arxiv.org/abs/2504.09689
Related Insights
Automated Safety Testing and Reporting Application for Conversational Safety Monitoring of Generative AI Tools for Mental Health: Development and Validation Study
JMIR Mental Health · 19 May 2026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
arXiv preprint · 30 Apr 2026
Between Help and Harm: An Evaluation Study of Mental Health Crisis Handling by Large Language Models
JMIR Mental Health · 11 Jun 2026