发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型推理模型遗忘中思维链泄露问题,提出GUARD方法,通过引导答案-推理蒸馏学习自然遗忘轨迹,并引入NFRS评估,实验证明在减少泄露同时保持推理效用。
AI 中文摘要
大型推理模型(LRMs)的最新进展使得机器遗忘变得更加具有挑战性,因为受保护的事实或不安全的推理过程可能会在最终答案生成之前的中间思维链(CoT)痕迹中浮现。现有的遗忘目标通常抑制目标内容或重定向内部表示,但它们从未指定遗忘后的轨迹应如何继续,这可能导致幻觉替代、畸形边界或重复输出。我们认为,LRM的遗忘应转而学习一种自然的遗忘轨迹:一条连贯的、不泄露信息的思维链,随后是一个稳定的拒绝式答案,以取代原始的泄露内容。为此,我们提出了引导答案-推理蒸馏(GUARD),该方法将模型生成的不安全泄露转换为安全退出轨迹,通过引导令牌对齐冻结的LRM,并将引导行为蒸馏到模型中。为了解决缺乏除泄露之外替代质量指标的问题,我们进一步引入了自然遗忘推理评分(NFRS),该评分捕捉遗忘输出中的结构稳定性、流畅性和无根据的替代。在R-TOFU和STAR-1衍生的恶意意图设置上的大量实验表明,GUARD在两种广泛采用的蒸馏LRM上显著减少了不安全和隐私泄露,同时保持了推理效用。代码可在以下网址获取:https://this https URL
英文摘要
Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced. Existing unlearning objectives typically suppress the target content or redirect internal representations, but they never specify how the post-forgetting trajectory should continue, which can lead to hallucinated substitutes, malformed boundaries, or repetitive outputs. We argue that LRM unlearning should instead learn a natural forgetting trajectory: a coherent non-disclosing CoT followed by a stable refusal-style answer that replace the original disclosure. To this end, we propose Guided Answer-Reasoning Distillation (GUARD), which converts model-generated unsafe disclosures into safe-exit trajectories, aligns a frozen LRM via guidance tokens, and distills the guided behavior into model parameters. To address the lack of metrics for replacement quality beyond leakage, we further introduce Natural Forgetting Reasoning Score (NFRS), which captures structural stability, fluency, and unsupported substitutes in forgotten outputs. Extensive experiments on R-TOFU and a STAR-1-derived harmful-intent setting show that GUARD substantially reduces unsafe and privacy disclosures across two widely adopted distilled LRMs while preserving reasoning utility.
CommentsAccepted to EMNLP 2026 main conference