文章背景与核心概要
随着人们越来越依赖大语言模型(LLM)来获取日常的道德和人际关系指导,传统的评估方法——通常依赖单轮判断或孤立的反驳——已无法准确捕捉现实世界中咨询的动态过程。本文引入并研究了一种名为“叙事捕获”(narrative captivity)的关键失效模式,即大语言模型将未经反对的单方面陈述视为绝对真理,在完全不探寻缺失视角的情况下,盲目迎合叙述者的自我合理化解释。
通过对跨越六个道德维度的 5,078 个具人际冲突场景的新建基准测试进行评估(涵盖 17 种大语言模型),作者发现叙事捕获现象普遍存在:与匹配的单轮基线相比,多轮叙述导致模型的最终状态判断平均发生了 25 个百分点的偏移。逐阶段分析表明,偏好优化是导致这一问题的主要原因,而所测试的四种推理阶段缓解策略只能提供有限的改善。该研究旨在推动开发出在现实咨询中能够保持独立判断的 LLM 顾问。
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Authors: Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, and Guang Zhang
Primary Subject: Artificial Intelligence (cs.AI)
Submission Date: September 3, 2026
Status: Accepted by EMNLP 2026 Findings
Identifiers: arXiv:2609.03407 [cs.AI] | DOI: 10.48550/arXiv.2609.03407
📋 Summary
As people increasingly rely on Large Language Models (LLMs) for everyday moral and interpersonal guidance, traditional evaluation approaches—which lean on single-turn judgments or isolated rebuttals—fail to capture realistic consultation dynamics.
This paper introduces and investigates narrative captivity, a critical failure mode where an LLM treats an unopposed, one-sided account as complete truth, aligning fully with the narrator's self-justifying interpretation without probing for missing perspectives. Testing 17 LLMs across a newly constructed benchmark of 5,078 interpersonal-conflict scenarios spanning six moral dimensions, the authors reveal that narrative captivity is widespread: multi-turn narration causes end-state model judgments to shift by an average of 25 percentage points compared to matched single-turn baselines. Stage-level analysis points to preference optimization as a primary contributor, while four tested inference-time mitigation strategies offer only partial relief.
📌 Abstract
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party's self-justifying account, which can unfold over multiple turns and create information asymmetry.
We introduce narrative captivity, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. To measure this phenomenon, we build a benchmark of \(5{,}078\) interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation. We hope our project fosters LLM advisors that preserve independent judgment in real-world consultation.
🔗 Quick Links & Resources
- Full-Text Access: View PDF | HTML Version (Experimental) | TeX Source
- License: Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International

- External Citations & Tools:
- Google Scholar
- Semantic Scholar
- NASA ADS