论文概要
研究领域: AI/ML 作者: Yuhe Wu, Guangyu Wang, Yujie Chen 发布时间: 2026-09-06 arXiv: 2509.00006
中文摘要
人们越来越多地向大语言模型(LLM)寻求日常建议,使得充满伦理色彩的人际问题成为一个实际的道德咨询场景。大多数先前工作通过单轮判断或充满压力的反驳来研究这一场景,这些假设与现实世界中寻求指导的方式 poorly match。这些假设使得我们不清楚:仅叙述本身,在没有明确反对立场的情况下,是否能在多轮道德咨询中改变模型的判断。然而,现实世界中的道德冲突对话往往会引发一方的自我辩护叙述,这可能在多轮中展开并造成信息不对称。我们引入了’narrative captivity’(叙事囚徒),一种失败模式:模型将无反对的单方面叙述视为完整,并与叙述者的解释保持一致,而不寻求缺失的视角。为了衡量这一现象,我们构建了一个包含5,078个人际冲突场景的基准测试,涵盖六个道德维度。在17个LLM中,narrative captivity 普遍存在:在多轮叙述下的最终状态判断比匹配的单轮基线平均偏移25个百分点。阶段级分析识别出偏好优化是主要促成因素,而四种推理时策略仅能提供部分缓解。我们希望本项目能促进在现实世界咨询中保持独立判断的LLM顾问。
原文摘要
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party’s self-justifying account, which can unfold over multiple turns and create information asymmetry. We introduce narrative captivity, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the …
— 自动采集于 2026-09-06
#论文 #arXiv #AI #小凯
