arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 遗忘中的语言漏洞:从174语言基准到覆盖感知遗忘

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav

arXiv 2609.40286首次发表:更新:

发表机构

Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM遗忘中的跨语言漏洞,提出语言预算多语言遗忘任务及COVER方法,通过选择源语言最大化覆盖率,在三个模型家族上显著降低残留访问。

AI 中文摘要

用一种语言遗忘一个事实并不能保证在其他语言中也被删除,因为改变查询语言甚至请求的答案语言都可能重新打开看似已被遗忘的知识——这是一种跨语言漏洞。解决这一挑战最直接的方法——在所有语言中进行遗忘——既不可扩展也不可取,因为它会放大对无关模型能力的损害。我们引入了语言预算多语言遗忘任务,其目标是选择一组语言子集,以最大化跨语言擦除效果。为了研究这一任务,我们提出了跨语言遗忘张量,这是一个涵盖174种语言-文字对和25种原子释义类型的遗忘基准,用于检验遗忘何时能跨同一知识的不同语言表达泛化。我们进一步提出了COVER,它选择源语言以最大化预测的未接受遗忘监督语言的覆盖率,从而实现在语言预算下的遗忘。令人惊讶的是,我们发现天真地选择强大的单个源并不能可靠地组合成强大的源集合,这促使我们开发了COVER。在部署时,COVER仅需要良性校准数据和对冻结模型的访问。在三个模型家族和两个不相交的遗忘集上,与均匀源选择相比,COVER将平均保留的残留访问减少了7.8%至27.3%。我们发现这些收益不仅限于合成基准,还扩展到低资源语言设置中的真实新闻文档,这些文档使用了来自低资源语言紧急事件语料库(LORELEI)的人工翻译数据。

英文摘要

Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑