发表机构
Northwestern University; New York University; University of Florida(西北大学; 纽约大学; 佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对上下文知识内部的冲突,构建了ContextConflict数据集,分析9种LLM的冲突处理能力,揭示其偏向早期证据的偏差,并提出无训练无标签的引导方法提升冲突解决效果。
AI 中文摘要
大多数现有研究聚焦于大语言模型(LLM)内部参数化知识与外部提供的上下文之间的冲突。与之不同,我们研究LLM如何处理上下文知识内部产生的冲突。我们提出了六种上下文冲突的分类(事实类、推理类、时间类、粒度类、视角类和歧义类),并为此场景构建了综合数据集ContextConflict。该数据集包含5781个样本,覆盖推理和摘要任务,包含显性矛盾及需要多步推理的隐性冲突。对9种LLM的实验表明,当前模型在解决上下文知识冲突方面仍存在不足。我们进一步对LLM处理此类冲突的过程进行了机械可解释性分析,揭示了其对冲突的潜在感知以及冲突处理背后的表征几何结构。此外,我们的分析还发现模型存在偏向早期证据的一致偏差,这种位置偏好是有效解决冲突的关键障碍。基于这些发现,我们提出了一种简单的无训练、无标签引导方法,该方法通过引导激活来鼓励更全面地整合证据,以实现更好的冲突解决。在我们的数据集上,该方法持续提升推理任务的准确率,并为摘要任务生成更高质量、更平衡的摘要。
英文摘要
Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual conflicts (factual, inferential, temporal, granularity, perspective, and ambiguity) and contribute a comprehensive dataset ContextConflict for this setting. The dataset contains 5,781 samples, covers both reasoning and summarization tasks, and includes both explicit contradictions and implicit conflicts that require multi-step reasoning. Experiments on nine LLMs show that current models still fall short in resolving contextual knowledge conflicts. We further provide mechanistic interpretability insights into how LLMs process such conflicts, revealing their latent awareness of conflicts and the representational geometry underlying conflict processing. In addition, our analysis uncovers a consistent model bias towards earlier evidence, and this positional preference serves as a key obstacle to effective conflict resolution. Motivated by these findings, we further propose a simple training-free, label-free steering method that steers activations to encourage a more comprehensive incorporation of evidences for better conflict resolution. On our dataset, the method consistently improves accuracy on reasoning tasks and generates higher-quality, more balanced summaries for summarization tasks.
CommentsAccepted to EMNLP 2026