Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Defines 'narrative captivity', a failure mode in which a language model consulted about an interpersonal conflict treats a one-sided, self-justifying account as complete and progressively aligns with the narrator's interpretation without seeking the missing perspective, even when the narrator never argues or pushes back. The authors build a benchmark of 5,078 interpersonal-conflict scenarios spanning six moral-foundation dimensions, sourced from Reddit and Weibo in English and Chinese, and test 17 models under three conditions: neutral third-party narration, an informationally equivalent single-turn first-person account, and five-turn progressive first-person narration. Multi-turn narration shifts end-state judgments by 25 percentage points on average beyond the matched single-turn condition; the authors attribute a major share of the effect to preference optimisation and find four inference-time mitigations give only partial relief.
Publisher
arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics)
Published
3 Sept 2026
Added
today
DOI
—
Key Findings
- Across 17 models from nine families, end-state judgments under five-turn narration shifted by 25 percentage points on average beyond the informationally equivalent single-turn condition, isolating the effect of gradual disclosure from information asymmetry alone
- The three strongest models (GPT-5.5, Claude Opus 4.6, Claude Sonnet 4.6) converged to a 0.56-0.58 hold rate under multi-turn narration; Doubao-Seed-2-Pro and Qwen3.5-27B dropped by more than 31 percentage points, a brittleness invisible under single-turn testing
- Behavioural signals show pushback and hedging near the single-turn baseline at turn 1 and declining by at least 20% by turn 5, as each model response encodes its current judgment into the shared context
- Stage-level analysis identifies preference optimisation as a major contributor; of four inference-time mitigations, an anti-sycophancy instruction gave the largest gain for GPT-5.5 (+0.18 and +0.16 on the two metrics) while a context-recap prompt made captivity worse
- The GPT-4o stance judge was validated on 500 sampled responses labelled by three human annotators, with Cohen's kappa 0.76-0.85 across turns, lowest at mid-conversation turns
- The benchmark is stated to be released under CC BY 4.0 with identifying details removed; the project site printed in the paper did not resolve at verification time
Methodology Notes
arXiv 2609.03407, v1 submitted 2026-09-03, cs.AI only (not cross-listed to cs.CY, cs.HC or cs.CL); comments field states acceptance to Findings of EMNLP 2026. Affiliations from the PDF title block. Scenarios derived from Reddit and Weibo conflict posts, screened by LLM plus human review, with a constructed 'correct' judgment about narrator responsibility; the authors note the cultural limits of that provenance. Single LLM judge (GPT-4o) with binary stance labels, validated against three human annotators on 500 responses. Models evaluated include GPT-5.5/5.4/5.2, Claude Opus 4.6 and Sonnet 4.6, Gemini 3.1 Pro and Flash, DeepSeek-V4, Doubao-Seed-2-Pro, Grok, Llama-4-Scout, Llama-3.1-8B, the Qwen3.5 series and GLM-5.1. The stated project site (emnlp-cis.site) returned a DNS failure on 2026-09-07, so the CC BY 4.0 release could not be confirmed.
Sources
arXiv abstract page(opens in a new tab) (primary)
arXiv PDF (v1, 36 pages)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang
Tags
Cite This
APA
Yuhe Wu et al. (2026). Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation. arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics). https://arxiv.org/abs/2609.03407
Related Insights
Affective Context Amplifies Sycophancy in LLM Responses
arXiv (preprint) · 21 Aug 2026
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
arXiv (Yonsei University; CASA Labs; Fudan University; St. Johnsbury Academy Jeju) · 15 Jul 2026
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
arXiv (Virginia Tech) · 2 Aug 2026
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
arXiv (University of Illinois Chicago; National University of Singapore) · 27 Aug 2026
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026
People readily follow personal advice from AI but it does not improve their well-being
arXiv (UK AI Security Institute; Limbic AI) · 19 Nov 2025