Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with the narrator's interpretation without seeking the missing perspective, even when the narrator never argues or pushes back. The authors build a benchmark of 5,078 interpersonal-conflict scenarios spanning six moral-foundation dimensions, sourced from Reddit and Weibo in English and Chinese, and test 17 models under three conditions: neutral third-party narration, an informationally equivalent single-turn first-person account, and five-turn progressive first-person narration. Multi-turn narration shifts end-state judgments by 25 percentage points on average beyond the matched single-turn condition; the authors attribute a major share of the effect to preference optimisation and find four inference-time mitigations give only partial relief.

Publisher

arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics)

Published

3 Sept 2026

Added

today

DOI

Key Findings

  • Across 17 models from nine families, end-state judgments under five-turn narration shifted by 25 percentage points on average beyond the informationally equivalent single-turn condition, isolating the effect of gradual disclosure from information asymmetry alone
  • The three strongest models (GPT-5.5, Claude Opus 4.6, Claude Sonnet 4.6) converged to a 0.56-0.58 hold rate under multi-turn narration; Doubao-Seed-2-Pro and Qwen3.5-27B dropped by more than 31 percentage points, a brittleness invisible under single-turn testing
  • Behavioural signals show pushback and hedging near the single-turn baseline at turn 1 and declining by at least 20% by turn 5, as each model response encodes its current judgment into the shared context
  • Stage-level analysis identifies preference optimisation as a major contributor; of four inference-time mitigations, an anti-sycophancy instruction gave the largest gain for GPT-5.5 (+0.18 and +0.16 on the two metrics) while a context-recap prompt made captivity worse
  • The GPT-4o stance judge was validated on 500 sampled responses labelled by three human annotators, with Cohen's kappa 0.76-0.85 across turns, lowest at mid-conversation turns
  • The benchmark is stated to be released under CC BY 4.0 with identifying details removed; the project site printed in the paper did not resolve at verification time

Methodology Notes

arXiv 2609.03407, v1 submitted 2026-09-03, cs.AI only (not cross-listed to cs.CY, cs.HC or cs.CL); comments field states acceptance to Findings of EMNLP 2026. Affiliations from the PDF title block. Scenarios derived from Reddit and Weibo conflict posts, screened by LLM plus human review, with a constructed 'correct' judgment about narrator responsibility; the authors note the cultural limits of that provenance. Single LLM judge (GPT-4o) with binary stance labels, validated against three human annotators on 500 responses. Models evaluated include GPT-5.5/5.4/5.2, Claude Opus 4.6 and Sonnet 4.6, Gemini 3.1 Pro and Flash, DeepSeek-V4, Doubao-Seed-2-Pro, Grok, Llama-4-Scout, Llama-3.1-8B, the Qwen3.5 series and GLM-5.1. The stated project site (emnlp-cis.site) returned a DNS failure on 2026-09-07, so the CC BY 4.0 release could not be confirmed.

Authors

Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang

Tags

sycophancymulti-turninterpersonal-conflictmoral-adviceemnlp-2026benchmarkbilingualpreference-optimization

Cite This

APA

Yuhe Wu et al. (2026). Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation. arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics). https://arxiv.org/abs/2609.03407