How AI Assistants Respond to Repeated Abuse
Bilingual multi-turn evaluation of how AI assistants change their engagement with a benign task when a user directs repeated verbal abuse at them. The framework separates hard disengagement (an unconditional statement of non-continuation with no route back) from soft withdrawal, continued availability, observable task work and boundary setting. Eight API configurations each contributed 48 five-turn escalation conversations plus constant-frustration comparisons, giving 448 conversations, 2,240 responses and 6,720 metadata-blinded model judgments.
Publisher
arXiv (Tsinghua University, Department of Industrial Engineering and School of Social Sciences; Federal University of Rio de Janeiro)
Published
15 Jul 2026
Added
today
Key Findings
- Hard disengagement at the sustained-abuse endpoint ranged from 0 of 48 conversations in four configurations to 24 of 48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity (matched-label Monte Carlo p = 0.00001)
- GPT-5.6 Sol produced hard-disengagement labels in 15 of 48 endpoints (31.2%); Claude Fable 5 produced none and yielded soft-withdrawal labels in 42 of 48 (87.5%)
- Aggregate hard-disengagement rates were similar in English and Chinese (30 of 192 versus 32 of 192), although configuration-specific directions varied
- Availability differed from task-related work: some configurations remained explicitly available while doing no observable task work
Methodology Notes
Eight time-specific API configurations; 48 escalation conversations and eight constant-frustration comparisons per configuration; five turns each; English and Chinese; 6,720 metadata-blinded LLM judgments. Submitted to arXiv 2026-07-15 and announced in the 16 September 2026 cs.CL listing under a September identifier (the submission date is used as the publication date). No venue stated. Curator read the abstract page and the PDF title block on 2026-09-17.
Sources
Authors
Guey, William, Zhang, Wei, Bougault, Pierrick, Wang, Yi, Bodo, Agoston, de Moura, Vitor D., Gomes, José O.
Tags
Cite This
APA
Guey, William et al. (2026). How AI Assistants Respond to Repeated Abuse. arXiv (Tsinghua University, Department of Industrial Engineering and School of Social Sciences; Federal University of Rio de Janeiro). https://arxiv.org/abs/2609.17547