arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08067cs.CL

流行知识在LLM知识更新中传播更多错误

Popular Knowledge Propagates More Errors in LLM Knowledge Updating

Yuji Zhang, Weibing Wang, Cheng Qian, Duo Zhou, Dilek Hakkani-Tür, Kathleen McKeown, Chengxiang Zhai, Heng Ji

首次发表
浏览论文内容

中文总结 AI 辅助

研究发现LLM中与高连接实体相关的流行事实在知识更新时更易被破坏并传播错误,据此提出轻量级重放策略PopAnchor以保留流行事实、减少遗忘。

中文摘要 AI 辅助

通过微调更新语言模型的知识对于保持其输出时效性至关重要,但同时也可能引发事实遗忘和新的幻觉。先前研究表明,长尾知识更难获取,且新记忆的长尾事实在后续微调中难以保留。我们研究一个互补的问题:在模型已正确编码的事实中,哪些事实在其它更新期间最易受到连带破坏?为了在现实的事实分布下探究此问题,我们构建了一个大规模图FACTPROP,包含经过验证的维基百科事实,通过连接共享头实体或尾实体的三元组,从而保留事实知识之间的关联。我们在事实陈述上微调模型,并在每次更新后测量从正确变为错误的事实。我们的结果揭示了一种与先前关于长尾脆弱性在获取和保留阶段发现不同的模式:在模型已能正确回答的事实中,那些与高度连接实体相关的事实更可能被邻近更新破坏,且对这些事实的更新会传播更广泛的错误。因此,结构流行度既能预测脆弱性,也能预测下游损害。受此发现启发,我们提出了基于流行度的锚定(PopAnchor),一种轻量级重放策略,保留少量流行事实并减少遗忘。

英文摘要

Updating a language model's knowledge through fine-tuning is essential for keeping its outputs current, yet can also induce factual forgetting and new hallucinations. Prior work shows that long-tail knowledge is harder to acquire and newly memorized long-tail facts are difficult to retain during later fine-tuning. We study a complementary question: among facts that a model has encoded correctly, which are most vulnerable to collateral corruption during other updates? To investigate this question under a realistic factual distribution, we construct a large-scale graph FACTPROP of verified Wikipedia facts by linking triples that share head or tail entities, thereby preserving connections among factual knowledge. We fine-tune models on factual statements and measure correct-to-incorrect facts after each update. Our results reveal a pattern distinct from prior findings on long-tail vulnerability during acquisition and retention: among facts that models already answer correctly, those associated with highly connected entities are more likely to be corrupted by neighboring updates, and updates to such facts propagate errors more broadly. Structural popularity therefore predicts both vulnerability and downstream damage. Inspired by this finding, we propose Popularity-based Anchoring (PopAnchor), a lightweight rehearsal strategy that preserves a small set of popular facts and reduces forgetting.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Columbia University(哥伦比亚大学)
  • City University of New York(纽约城市大学)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑