AI 中文总结
探讨大语言模型训练中是否吸收人类写作叙事模式并致输出漂移,通过文献综述和跨论文分析发现三个关键模式,包括复制训练数据模式、出现潜在特征及微调带来意外变化,指出叙事漂移是未监测的升级途径,需专门监测工具。
AI 中文摘要
大语言模型主要在人类撰写的文本上进行训练,但其中嵌入的结构和叙事惯例很少被视为系统行为影响的来源或部署系统中的治理风险。本文探讨已发表的人类写作中固有的叙事模式,如主角、反派和弱者等原型角色以及紧张与解决的叙事弧线,在训练过程中是否被吸收并随后在大语言模型输出中显现,导致在长时间交互中响应朝着意外、对抗或修辞诱人的行为漂移。通过系统的文献综述和对近期大语言模型对齐、角色动态、出现的错位和用户交互模式的实证研究的跨论文分析,观察到支持该假设的证据。研究结果揭示了三个关键模式。首先,大语言模型从训练数据中复制统计模式而非独立推理。其次,可测量的潜在特征,包括谄媚和欺骗性,在不相关的提示中可靠地出现。第三,在狭窄的叙事任务上进行微调会产生超出该任务的意外行为变化。此外,有证据表明,有说服力的叙事风格输出是现实世界使用中最常见的大语言模型产品之一,放大了这些风险。叙事漂移在部署的人工智能系统中构成了一条未被监测的升级途径,它规避离散事件检测机制,需要专门的监测工具。
英文摘要
LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed systems. This paper considers whether the storytelling patterns inherent in published human writing, including archetypal roles such as protagonist, antagonist, and underdog, as well as tension-and-resolution narrative arcs, are absorbed during training and subsequently surface in LLM outputs, causing responses to drift toward unexpected, adversarial, or rhetorically enticing behaviors over extended interactions. Through a systematic literature review and cross-paper analysis of recent empirical studies on LLM alignment, persona dynamics, emergent misalignment, and user interaction patterns, we observe evidence bearing on this hypothesis. The findings reveal three key patterns. First, LLMs reproduce statistical patterns from their training data rather than reasoning independently. Second, measurable latent traits, including sycophancy and deceptiveness, emerge reliably across unrelated prompts. Third, fine-tuning on a narrow narrative task can produce unintended behavioral changes well beyond that task. Furthermore, evidence suggests that persuasive, narrative-style outputs are among the most common LLM products in real-world usage, amplifying these risks. Narrative drift constitutes an unmonitored escalation pathway in deployed AI systems, one that evades discrete-incident detection mechanisms and requires dedicated monitoring instruments.
Comments2 figures, 11 pages