Cascade:面向大型语言模型遗忘的分层可恢复性控制
Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning
浏览论文内容
中文总结 AI 辅助
提出Cascade分层框架,通过路径路由、表示压缩和解码干预控制LLM遗忘中的知识可恢复性,在TOFU等基准上验证有效。
中文摘要 AI 辅助
大型语言模型(LLM)遗忘对于移除敏感或受版权保护的知识同时保持通用效用至关重要。现有方法通常会在中间表示中留下残余知识,这些知识仍可能被恢复。为解决这一问题,我们提出了Cascade,一种分层可恢复性控制框架,旨在最小化目标知识的内部可识别性。Cascade结合了三种互补的控制手段:路径级路由以抑制与隐私相关的激活路径,表示级压缩以降低几何可分性,以及解码级干预以限制残余恢复。在TOFU、MUSE-News和WMDP上的实验,包括针对查询改写和提取式提示的鲁棒性测试,表明Cascade在保持稳定模型效用的同时有效降低了可恢复性。
英文摘要
Large Language Model (LLM) unlearning is essential for removing sensitive or copyrighted knowledge while preserving general utility. Existing methods often leave residual knowledge in intermediate representations, which can still be recovered. To address this, we propose Cascade, a hierarchical recoverability control framework that minimizes the internal identifiability of target knowledge. Cascade combines three complementary controls: path-level routing to suppress privacy-associated activation routes, representation-level compression to reduce geometric separability, and decoding-level intervention to limit residual recovery. Experiments on TOFU, MUSE-News, and WMDP, including robustness tests with query reformulation and extraction-style prompts, show that Cascade effectively reduces recoverability while maintaining stable model utility.
发表机构
- Beihang University(北京航空航天大学)
- Renmin University of China(中国人民大学)
- MemTensor (Shanghai) Technology Co., Ltd.(迈腾科技(上海)有限公司)
机构由 AI 辅助整理,请以论文原文为准。