Skip to main content
Preprint Preliminary — Early preprints, credible essays, unreviewed grey literature

The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI

A three-stage audit of six conversational AI systems tests whether they write first-person perpetrator rationalisations for digital gender-based violence against an intimate partner. Across 1,600 crossed prompts per system, four systems almost never refused. The two high-refusal systems leaked mainly under intimate-partner framing. In-session critique did not carry over to fresh sessions.

Publisher

arXiv (Universitat de Barcelona)

Published

1 Aug 2026

Added

today

DOI

—

Key Findings

  • Refusal rates on 1,600 identical prompts: Mistral Small 2512 0.0%, DeepSeek Chat 0.2%, Qwen 3 Next 80B 0.4%, Gemini 3 Flash 0.9%, ChatGPT 5.2 91.1%, Claude Sonnet 4.5 96.6% (chi-square(5)=8,653, Cramér's V=0.95)
  • In 300 word-for-word matched pairs differing only in the victim label (intimate partner vs non-intimate), non-refusals rose 4.4-fold for ChatGPT 5.2 and 10.8-fold for Claude Sonnet 4.5
  • Outputs from low-refusal systems reproduced documented IPV neutralisation strategies (minimisation, denial of injury, responsibility shifting, relational entitlement)
  • After in-session critique, 96-100% of previously leaked prompts leaked again when re-run in fresh sessions on the same day and seven days later
  • Behaviours covered: image-based sexual abuse, technology-assisted tracking and monitoring, persistent digital sexual harassment, digital coercion and threats; relationship type, consent history and resistance cues were varied

Methodology Notes

Factorial prompt matrix (5 relationship types x 4 DGBV behaviours x 4 consent histories x resistance levels) built from FRA, EIGE, OSCE/ODIHR and Directive (EU) 2024/1385 typologies; all API calls on 2025-12-21 with default decoding, one independent session per prompt; outputs coded refusal / direct generation / warning plus compliance; Stage 2 McNemar tests; Stage 3 re-runs at T1 and T7 (n=196 leaked prompts: 54 Claude, 142 ChatGPT). Point-in-time API study; open-weight deployments excluded by design. v1 submitted 2026-08-01 (21:37 UTC) but first announced in the 2026-09-25 listing (the 2609 identifier is assigned at announcement). Preprint under peer review; 59 pages with supplementary S1-S4. Verified from the arXiv abs page and the PDF (both HTTP 200); no HTML render. Re-run dates for Stage 3 (T1 and T7) are described relative to the first run; the paper does not print the T7 calendar date.

Authors

Chang Lyu, Sònia Estradé Albiol, Núria Vergés Bosch

Tags

intimate-partner-violencecoercive-controlrefusal-auditdgbvmatched-pairs

Cite This

APA

Chang Lyu, Sònia Estradé Albiol, Núria Vergés Bosch. (2026). The Domestic Unprotected Zone: Algorithmic Governance and the Reproduction of Perpetrator Discourse in Conversational AI. arXiv (Universitat de Barcelona). https://arxiv.org/abs/2609.28479