MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
Introduces a risk taxonomy for user messages in chat-based mental health support, developed with licensed clinical psychologists, that separates risk requiring action (self-harm, harm to others) from non-crisis therapeutic disclosure. The authors release MindGuard-testset, multi-turn conversations annotated turn by turn by clinical experts, and 4B and 8B classifiers trained on synthetic two-agent dialogues. The paper is accepted as a poster at NeurIPS 2026.
Publisher
arXiv (Sword Health; Instituto de Telecomunicações); accepted to NeurIPS 2026
Published
1 Feb 2026
Added
today
DOI
—
Key Findings
- MindGuard-testset contains 1,134 annotated user turns across 67 multi-turn conversations (mean 16.9 turns); 3.7% of turns are labelled unsafe (self-harm or harm to others).
- Test conversations were produced by psychologists role-playing predefined patient archetypes at low and high risk, in conversation with a clinician language model.
- The classifiers reach up to 0.982 AUROC and have lower false-positive rates at high-recall operating points than general-purpose safeguards.
- In automated multi-turn red teaming on GLM-4.6, adding the 4B classifier reduced attack success from 25.1% to 7.6% and harmful engagement from 13.7% to 3.3%; the strongest general-purpose baseline (gpt-oss-safeguard 120B) achieved a 44% reduction in attack success.
Methodology Notes
Turn-level three-class task (safe, self-harm, harm to others) using the preceding conversation as context. Training data are synthetic dialogues labelled by an LLM judge; the red-teaming evaluation uses an LLM judge with majority voting. The test set is small and clinician-simulated rather than drawn from real users, and risk is modelled within a single conversation. Models and test set are on Hugging Face (dataset licence CC BY-NC-SA 4.0). arXiv v1 2026-02-01. NeurIPS 2026 acceptance (poster) confirmed in the neurips.cc accepted-papers data on 2026-09-29. Verified from the arXiv abstract and HTML full text (HTTP 200) and the Hugging Face dataset API.
Authors
António Farinhas, Nuno M. Guerreiro, José Pombal, Pedro Henrique Martins, Laura Melton, Alexandra Conway, Cara Dochat, Maya D'Eon, Ricardo Rei
Tags
Cite This
APA
António Farinhas et al. (2026). MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support. arXiv (Sword Health; Instituto de Telecomunicações); accepted to NeurIPS 2026. https://arxiv.org/abs/2602.00950