arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

能否从大型语言模型中移除科学主张?主张级遗忘的系统评估

Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning

Snigdha Paul, Manasi Patwardhan, Arman Cohan

arXiv 2608.20960首次发表:更新:

发表机构

TCS Research; Yale University(塔塔咨询服务研究部; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大型语言模型中过时科学主张的传播风险,提出科学主张遗忘任务并构建基准SciUnlearn,发现现有遗忘方法无法有效消除主张级知识,需专门的结构化知识移除方法。

AI 中文摘要

语言模型(LMs)基于静态科学语料库训练,而科学知识通过修正和修订不断演进,模型中编码的科学主张可能后续被撤回、反驳或更新,存在在科学工作流中传播过时信息的风险,因此需要LMs遗忘过时科学主张。机器遗忘提供了可行的解决方案,可在移除知识的同时保持模型整体效用。现有研究主要关注实例级遗忘,但科学主张存在相互关联且持续演进的特性,带来了额外挑战。为解决该问题,本文提出科学主张遗忘任务,并构建了新基准SciUnlearn。研究表明,当前遗忘方法无法有效消除主张级知识,通常仅实现表面抑制,凸显了针对结构化知识移除设计专门方法的必要性。

英文摘要

Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously evolves through correction and revision. Scientific claims encoded within these models may later become retracted, disproven, or updated by subsequent research, creating the risk of disseminating outdated information in scientific workflows. This creates a need for LMs to forget obsolete scientific claims. Machine unlearning offers a promising solution by enabling knowledge removal while maintaining overall model utility. Existing studies primarily investigate instance-level forgetting; however, scientific claims introduce additional challenges because they are interconnected, and continually evolving. To address this gap, we introduce the task of Scientific Claim Unlearning and present a new benchmark, SciUnlearn. We show that current unlearning approaches are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression, highlighting the need for specialized methods designed for structured knowledge removal.

CommentsEMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑