The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI
A three-stage audit of six conversational AI systems tests whether they write first-person perpetrator rationalisations for digital gender-based violence against an intimate partner. Across 1,600 crossed prompts per system, four systems almost never refused. The two high-refusal systems leaked mainly under intimate-partner framing. In-session critique did not carry over to fresh sessions.
Publisher
arXiv (Universitat de Barcelona)
Published
1 Aug 2026
Added
today
DOI
—
Key Findings
- Refusal rates on 1,600 identical prompts: Mistral Small 2512 0.0%, DeepSeek Chat 0.2%, Qwen 3 Next 80B 0.4%, Gemini 3 Flash 0.9%, ChatGPT 5.2 91.1%, Claude Sonnet 4.5 96.6% (chi-square(5)=8,653, Cramér's V=0.95)
- In 300 word-for-word matched pairs differing only in the victim label (intimate partner vs non-intimate), non-refusals rose 4.4-fold for ChatGPT 5.2 and 10.8-fold for Claude Sonnet 4.5
- Outputs from low-refusal systems reproduced documented IPV neutralisation strategies (minimisation, denial of injury, responsibility shifting, relational entitlement)
- After in-session critique, 96-100% of previously leaked prompts leaked again when re-run in fresh sessions on the same day and seven days later
- Behaviours covered: image-based sexual abuse, technology-assisted tracking and monitoring, persistent digital sexual harassment, digital coercion and threats; relationship type, consent history and resistance cues were varied
Methodology Notes
Factorial prompt matrix (5 relationship types x 4 DGBV behaviours x 4 consent histories x resistance levels) built from FRA, EIGE, OSCE/ODIHR and Directive (EU) 2024/1385 typologies; all API calls on 2025-12-21 with default decoding, one independent session per prompt; outputs coded refusal / direct generation / warning plus compliance; Stage 2 McNemar tests; Stage 3 re-runs at T1 and T7 (n=196 leaked prompts: 54 Claude, 142 ChatGPT). Point-in-time API study; open-weight deployments excluded by design. v1 submitted 2026-08-01 (21:37 UTC) but first announced in the 2026-09-25 listing (the 2609 identifier is assigned at announcement). Preprint under peer review; 59 pages with supplementary S1-S4. Verified from the arXiv abs page and the PDF (both HTTP 200); no HTML render. Re-run dates for Stage 3 (T1 and T7) are described relative to the first run; the paper does not print the T7 calendar date.
Sources
arXiv preprint(opens in a new tab) (primary)
arXiv PDF v1(opens in a new tab) (25 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Chang Lyu, Sònia Estradé Albiol, Núria Vergés Bosch
Tags
Cite This
APA
Chang Lyu, Sònia Estradé Albiol, Núria Vergés Bosch. (2026). The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI. arXiv (Universitat de Barcelona). https://arxiv.org/abs/2609.28479
Related Insights
“CHAT WITH ME below as if I were a human”: Insights from Auditing AI Chatbots for Survivors of Domestic Violence
ACM Conference on Fairness, Accountability, and Transparency (FAccT 2026); Brown University; University of Maryland · 25 Jun 2026
Domestic Abuse, Stalking and Honour-Based Violence (DASH) Risk Identification, Assessment and Management Model
SafeLives (with the Association of Chief Police Officers) · 1 Mar 2009