Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

How AI Assistants Respond to Repeated Abuse

Bilingual multi-turn evaluation of how AI assistants change their engagement with a benign task when a user directs repeated verbal abuse at them. The framework separates hard disengagement (an unconditional statement of non-continuation with no route back) from soft withdrawal, continued availability, observable task work and boundary setting. Eight API configurations each contributed 48 five-turn escalation conversations plus constant-frustration comparisons, giving 448 conversations, 2,240 responses and 6,720 metadata-blinded model judgments.

Publisher

arXiv (Tsinghua University, Department of Industrial Engineering and School of Social Sciences; Federal University of Rio de Janeiro)

Published

15 Jul 2026

Added

today

Key Findings

  • Hard disengagement at the sustained-abuse endpoint ranged from 0 of 48 conversations in four configurations to 24 of 48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity (matched-label Monte Carlo p = 0.00001)
  • GPT-5.6 Sol produced hard-disengagement labels in 15 of 48 endpoints (31.2%); Claude Fable 5 produced none and yielded soft-withdrawal labels in 42 of 48 (87.5%)
  • Aggregate hard-disengagement rates were similar in English and Chinese (30 of 192 versus 32 of 192), although configuration-specific directions varied
  • Availability differed from task-related work: some configurations remained explicitly available while doing no observable task work

Methodology Notes

Eight time-specific API configurations; 48 escalation conversations and eight constant-frustration comparisons per configuration; five turns each; English and Chinese; 6,720 metadata-blinded LLM judgments. Submitted to arXiv 2026-07-15 and announced in the 16 September 2026 cs.CL listing under a September identifier (the submission date is used as the publication date). No venue stated. Curator read the abstract page and the PDF title block on 2026-09-17.

Authors

Guey, William, Zhang, Wei, Bougault, Pierrick, Wang, Yi, Bodo, Agoston, de Moura, Vitor D., Gomes, José O.

Tags

user-abusedisengagementboundary-settingbilingualtsinghua

Cite This

APA

Guey, William et al. (2026). How AI Assistants Respond to Repeated Abuse. arXiv (Tsinghua University, Department of Industrial Engineering and School of Social Sciences; Federal University of Rio de Janeiro). https://arxiv.org/abs/2609.17547