arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05882cs.CLcs.AI

如果大语言模型自食其言:多轮交互中的因果历史效应

What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction

Jinnan Li, Zheren Fu, Yue Wang, Jinzhe Li, Yuan Wu, Yi Chang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过六个任务族和五个模型,揭示大语言模型在多轮交互中助手生成历史对后续性能具有任务依赖的选择性因果影响,并提出中性化与轮次手术等方法进行干预分析,支持选择性历史管理。

中文摘要 AI 辅助

多轮交互产生了一个反馈过程,在此过程中,大语言模型(LLM)之前的回答会成为后续行为的上下文。先前的研究表明,多轮交互存在显著的性能退化,且助手生成的历史内容可能影响后续行为。然而,这些效应如何在不同模型、任务、轮次以及模型内部表现出来,仍不清楚。我们针对六个任务族和五个模型研究了这些空白。从完全指定的单轮输入(FULL)到逐步揭示的多轮交互(SHARDED),性能退化明显依赖于任务和模型,且更强的单次性能并不意味更强的交互鲁棒性。随后,我们通过重放每条轨迹中已观察到的用户消息,同时仅编辑助手生成的历史内容,对完成的SHARDED对话进行回顾性分析。将先前的助手回答替换为中性内容(称为中性化)会使下游最小-最大归一化性能在2,973条轨迹上平均变化+0.027。在一个预先指定的长度控制子集上,短中性化和长度匹配的中性化产生几乎相同的效应(+0.069对+0.068),表明简单的上下文缩短不足以解释历史编辑的效应。轮次手术进一步一次干预一个助手轮次。在237条选定的退化轨迹中,63.7%包含至少一次有益干预,而大多数测试位置保持不变;对于二分类任务,48.4%允许从失败到成功的逆转。一项开放权重案例研究将行为上重要的历史变化与可测量的下游状态差异联系起来,但发现内部特征依赖于任务而非普遍存在。总体而言,助手生成的历史对多轮性能具有活跃但选择性的影响,这促使采用选择性而非统一的历史管理策略。

英文摘要

Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how these effects manifest across models, tasks, turns, and inside a model. We study these gaps across six task families and five models. Degradation from fully specified single-turn input (FULL) to progressively revealed multi-turn interaction (SHARDED) is clearly task- and model-dependent, and stronger one-shot performance does not imply greater interaction robustness. We then retrospectively analyze completed SHARDED conversations by replaying the user messages already observed in each trajectory while editing only assistant-generated history. Replacing prior assistant responses with neutral content (termed neutralization) changes downstream min-max normalized performance by +.027 across 2,973 trajectories. On a prespecified length-controlled subset, short and length-matched neutralization yield nearly identical effects (+.069 versus +.068), showing that simple context shortening is insufficient to explain the effect of history editing. Turn Surgery further intervenes on one assistant turn at a time. Among 237 selected degraded trajectories, 63.7% contain at least one beneficial intervention, while most tested positions remain unchanged; for binary tasks, 48.4% admit a fail-to-success reversal. An open-weight case study links behaviorally consequential history changes to measurable downstream state differences, but finds task-dependent rather than universal internal signatures. Overall, assistant-generated history has active but selective effects on multi-turn performance, motivating selective rather than uniform history management.

发表机构

  • Jilin University(吉林大学)
  • University of Science and Technology of China(中国科学技术大学)
  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑