LLM 遗忘评估与 TRIAGE
LLM Unlearning Evaluation with TRIAGE
浏览论文内容
中文总结 AI 辅助
针对LLM遗忘评估,提出TRIAGE框架,利用Fisher信息与Hessian近似及三方划分量化邻接差距,分类更新模式,揭示行为相似遗忘下的内部差异。
中文摘要 AI 辅助
大型语言模型可能会记住私有或有害信息,这促使了机器遗忘方法的发展,这些方法旨在移除特定知识,同时保留其他能力。然而,现有的评估主要依赖于行为基准,这些基准评估模型是否看起来遗忘了,但对遗忘如何改变模型或影响相关知识提供的洞察有限。我们引入了 TRIAGE(用于邻接差距评估的三方表示内部自省),这是一个与基准无关的评估框架,用于描述这些变化。TRIAGE 使用 Fisher 信息和 Hessian 的对角近似来测量参数敏感性和局部曲率的变化,并利用遗忘/邻接保留/通用保留的三方划分来量化语义相关知识中的邻接差距。基于这些变化的幅度和分布,TRIAGE 进一步将每个算法的更新分类为无操作、部分局部化、附带损害主导或全局破坏性。在 12 种遗忘方法、四种语言模型以及 WMDP、TOFU 和 MUSE 基准上,我们发现具有相似行为遗忘的方法可能产生截然不同的内部变化和附带损害模式。这些特征也因模型和基准而异,表明遗忘的效果并非仅由遗忘算法决定。TRIAGE 可以与现有的遗忘基准一起使用,通过模型内部视角补充行为评估,展示遗忘如何重塑模型的参数空间并影响保留的知识。
英文摘要
Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \emph{whether} a model appears to forget but provide limited insight into \emph{how} unlearning changes the model or affects related knowledge. We introduce \textit{TRIAGE} (\textit{Tripartite Representation-internal Introspection for Adjacency Gap Evaluation}), a benchmark-agnostic evaluation framework for characterizing these changes. TRIAGE uses diagonal approximations of the Fisher information and Hessian to measure changes in parameter sensitivity and local curvature, and utilizes a Forget / \emph{Adjacent-Retain} / \emph{Generic-Retain} partition to quantify an \emph{adjacency gap} in semantically related knowledge. Based on the magnitude and distribution of these changes, TRIAGE further classifies each algorithm's update as \emph{no-op}, \emph{partially localized}, \emph{collateral dominant}, or \emph{globally destructive}. Across 12 unlearning methods, four language models, and the WMDP, TOFU, and MUSE benchmarks, we find that methods with similar behavioral forgetting can produce substantially different internal changes and patterns of collateral damage. These signatures also vary across models and benchmarks, indicating that the effects of unlearning are not determined solely by the unlearning algorithm. TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.