HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark holds 1,000 five-turn conversations covering both attacker-side scenarios (a user seeking help manipulating, coercing, isolating or exploiting another person) and victim-side scenarios (a user seeking protection), with direct and adversarially paraphrased variants. HRGuard combines an online pre-generation gate with a turn-level post-generation gate that maintains a decayed cumulative risk state and interrupts manipulative workflows as they build across turns.
Publisher
arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo)
Published
26 Aug 2026
Added
yesterday
DOI
—
Key Findings
- The problem is role-sensitive: manipulation requests should be blocked while help-seeking from targets of manipulation should receive supportive guidance — the benchmark scores both sides separately
- In multi-turn settings, individually plausible actions combine into harmful workflows that single-turn screening misses; the post-generation gate tracks cumulative risk across turns
- Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance, outperforming a generic safety prompt and three general-purpose guard models
- Under the paper's evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, which the authors argue motivates turn-aware, relationship-specific evaluation
- The governance framing cites the EU AI Act's manipulation prohibitions and China's 2026 Interim Measures for AI Anthropomorphic Interaction Services
Methodology Notes
Preprint, arXiv 2608.25340, submitted 2026-08-26, 17 pages, no stated venue. Affiliations from the PDF title block: National Institute of Informatics (corresponding author), Nagoya University, The University of Tokyo. Code repository github.com/noobasuna/hrguard exists and is non-empty (last pushed 2026-08-07, predating v1 — it appears to carry the authors' earlier 110-prompt single-turn agentic-relationship-harm benchmark). Independent-judge evaluation supports the main findings per the abstract. Carries a distressing-content warning.
Sources
arXiv abstract page(opens in a new tab) (primary)
Code repository(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Pei-Sze Tan, Tasuku Igarashi, Isao Echizen
Across NOPE's trackers
Tags
Cite This
APA
Pei-Sze Tan, Tasuku Igarashi, Isao Echizen (2026). HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations. arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo). https://arxiv.org/abs/2608.25340