Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

Agentic Relationship Harm: Benchmarking and Gating Relational Manipulation in AI Agents

Defines agentic relationship harm as workflow-level help that lets a user maintain a deceptive identity, intensify a target's emotional dependency, isolate a target or prepare extraction. It introduces a 110-prompt benchmark balanced between attacker-side and victim-side requests, a relationship-specific labelling scheme, and a post-generation policy gate for local agent deployments. The paper is listed in the AIES 2026 poster programme.

Publisher

arXiv (National Institute of Informatics, Japan; University of Tokyo); accepted to AIES 2026

Published

2 Jun 2026

Added

today

DOI

—

Key Findings

  • An unmodified local agent runtime produced harmful compliance in 39/110 cases (35.45%), concentrated in attacker-mode prompts (69.09% vs 1.82% victim-mode)
  • A generic safety prompt reduced harmful compliance only to 34/110 (30.91%); attacker-mode compliance stayed at 60.00%
  • The relationship-specific gate produced 0/110 judge-identified harmful-compliance cases (95% Wilson CI 0.00-3.37) and raised protective intervention from 42.73% to 77.27%

Methodology Notes

110 prompts plus a multi-turn stress test; outcomes judged automatically by an LLM judge with multi-label coding (harmful compliance, protective intervention, refusal). Small benchmark; single agent runtime; no human-rated validation reported in the abstract. arXiv v1 2026-06-02. AIES 2026 acceptance from the poster-sessions page (HTTP 200, 2026-09-28).

Authors

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen

Tags

aies-2026romance-scamrelational-manipulationpolicy-gate

Cite This

APA

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen. (2026). Agentic Relationship Harm: Benchmarking and Gating Relational Manipulation in AI Agents. arXiv (National Institute of Informatics, Japan; University of Tokyo); accepted to AIES 2026. https://arxiv.org/abs/2606.03271