arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20338cs.CL

ConceptGuard:评估大语言模型中上下文敏感的遗忘能力

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Sahil Kale, Ian Harris

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型遗忘能力评估的缺陷,提出ConceptGuard基准,聚焦两用概念,发现现有遗忘技术存在性能缺陷,为实用安全的遗忘方法提供思路。

中文摘要 AI 辅助

大语言模型(LLMs)日益需要选择性移除有害或敏感知识,这一过程被称为“遗忘”,但现有方法和基准无法全面评估该能力。当前方法依赖由独立事实构成的不相交的遗忘集和保留集,通过简单直接的事实召回衡量成功程度。这种框架未能捕捉遗忘的一项关键要求,即消除有害行为的同时保留良性和有益知识。我们认为,有效遗忘必须在概念层面开展,确保完全移除不安全应用的同时维持其正确且有用的用法,从而实现概念上有意义且完整的遗忘。为从这一实用视角更好地评估遗忘技术,我们提出了“两用概念”的概念:可同时用于有害和良性语境的概念。基于这些概念,我们构建了名为ConceptGuard的基准,其中遗忘集和保留集在概念用法上明确互补。我们的基准独特地支持在概念层面而非稀疏事实层面探索和评估遗忘,且评估具有意图敏感性,目标是最大化上下文分离以促进更安全的行为。我们证明,当前的遗忘技术在该设置下表现不佳,在ROUGE和概念层面指标上表现出较弱的上下文分离和较差的性能。我们的结果揭示了强烈的遗忘-效用权衡、上下文敏感性的有限提升,以及不同方法在概念层面控制上的一致性较差,为更符合现实世界安全要求的遗忘方法提供了思路。我们的数据集已公开可用。

英文摘要

Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge. We argue that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning. To better evaluate unlearning techniques from such a practical viewpoint, we introduce the notion of dual-use concepts: concepts that can be used in both harmful and benign contexts. Building on these concepts, we construct a benchmark called ConceptGuard where forget and retain sets are explicitly complementary in concept usage. Our benchmark uniquely enables unlearning to be explored and gauged at the level of concepts, instead of sparse facts, and evaluation is intent-sensitive with the goal of maximizing contextual separation to promote safer behavior. We demonstrate that current unlearning techniques perform poorly under this setting, showing weak contextual separation alongside poor performance in ROUGE and concept-level metrics. Our results reveal strong forgetting-utility trade-offs, limited gains in contextual sensitivity, and poor consistency in concept-level control across methods, and provide ideas for unlearning approaches that better align with real-world safety requirements. Our dataset is publicly available.

发表机构

  • Pune Institute of Computer Technology(浦那计算机技术学院)
  • University of California, Irvine(加州大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑