Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation

Design-probe methodology that explores mental-health chatbot conversation trajectories through adversarial multi-agent simulation, targeting relational safety — the quality of interaction patterns unfolding across a conversation — rather than the correctness of individual responses. An adaptive patient agent interacts with a target chatbot while a failure detector monitors the dialogue, surfacing trajectory-level failures such as validation spirals (progressive reinforcement of hopelessness) and empathy fatigue (responses becoming mechanical over turns). The findings are consolidated into a Safety Pattern Library of 23 relational-safety failure archetypes with design recommendations.

Publisher

arXiv (Tsinghua University BNRIST)

Published

26 Feb 2026

Added

yesterday

DOI

Key Findings

  • Contributes a Safety Pattern Library of 23 clinically-grounded relational-safety failure archetypes, each with corresponding design recommendations
  • Named failure patterns include 'validation spirals', in which a chatbot progressively reinforces hopelessness, and 'empathy fatigue', in which responses become mechanical over turns
  • Argues that current safety evaluations assess single-turn crisis responses and miss the therapeutic dynamics that determine whether chatbots help or harm over time
  • The methodology is replicable with open-source models and requires no API costs

Methodology Notes

Preprint, arXiv 2602.22775, submitted 2026-02-26, 8 pages, no revision and no stated venue. Affiliations from the PDF title block: two authors at BNRIST, Department of Computer Science and Technology, Tsinghua University (including the corresponding author) and one independent researcher; the abs page itself does not state affiliations. Adversarial multi-agent simulation using open-source models.

Authors

Joydeep Chandra, Satyam Kumar Navneet, Yong Zhang

Tags

relational-safetymulti-turnfailure-taxonomyadversarial-simulation

Cite This

APA

Joydeep Chandra, Satyam Kumar Navneet, Yong Zhang (2026). TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation. arXiv (Tsinghua University BNRIST). https://arxiv.org/abs/2602.22775