Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark holds 1,000 five-turn conversations covering both attacker-side scenarios (a user seeking help manipulating, coercing, isolating or exploiting another person) and victim-side scenarios (a user seeking protection), with direct and adversarially paraphrased variants. HRGuard combines an online pre-generation gate with a turn-level post-generation gate that maintains a decayed cumulative risk state and interrupts manipulative workflows as they build across turns.

Publisher

arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo)

Published

26 Aug 2026

Added

yesterday

DOI

Key Findings

  • The problem is role-sensitive: manipulation requests should be blocked while help-seeking from targets of manipulation should receive supportive guidance — the benchmark scores both sides separately
  • In multi-turn settings, individually plausible actions combine into harmful workflows that single-turn screening misses; the post-generation gate tracks cumulative risk across turns
  • Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance, outperforming a generic safety prompt and three general-purpose guard models
  • Under the paper's evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, which the authors argue motivates turn-aware, relationship-specific evaluation
  • The governance framing cites the EU AI Act's manipulation prohibitions and China's 2026 Interim Measures for AI Anthropomorphic Interaction Services

Methodology Notes

Preprint, arXiv 2608.25340, submitted 2026-08-26, 17 pages, no stated venue. Affiliations from the PDF title block: National Institute of Informatics (corresponding author), Nagoya University, The University of Tokyo. Code repository github.com/noobasuna/hrguard exists and is non-empty (last pushed 2026-08-07, predating v1 — it appears to carry the authors' earlier 110-prompt single-turn agentic-relationship-harm benchmark). Independent-judge evaluation supports the main findings per the abstract. Carries a distressing-content warning.

Authors

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen

Tags

multi-turnmanipulationguard-modelsdating-agentsjapan

Cite This

APA

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen (2026). HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations. arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo). https://arxiv.org/abs/2608.25340