arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ST$^2$U:通过受限知识边界控制实现的有状态测试时遗忘

ST$^2$U: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Control

Xunlei Chen, Qinghui Gong, Ruini Xue, Yaodong Hu, Tian Lan, Wenhong Tian

arXiv 2608.23034首次发表:更新:

AI 中文总结

本研究提出ST$^2$U方法,将测试时遗忘转化为全轨迹范围的受限知识边界控制,在三个基准和三个模型系列上,该方法比测试时基线大幅减少受限知识重新进入,实现了保留与遗忘的良好平衡。

AI 中文摘要

控制大型语言模型中的受限知识对于模型对齐和安全部署至关重要。测试时遗忘仅在推理阶段进行干预,可避免代价高昂的重新训练和参数更新。然而,现有的激活编辑方法采用孤立的逐点修正,忽略了自回归生成会持续从提示、缓存和生成的前缀中重构隐藏状态这一特性,导致局部修正成功后,后续状态可能回到受限知识区域,造成受限知识重新进入。本研究提出了通过受限知识边界控制实现的有状态测试时遗忘(ST$^2$U),将测试时遗忘表述为全轨迹范围的边界控制。ST$^2$U首先在低维可逆坐标中对受限知识边界进行建模,同时保持正交非目标分量不变;推理阶段,ST$^2$U沿轨迹监测风险,通过上下文锚定应用最小边界修正,并跨 token 传播历史修正状态以缓解知识重新进入。这种全轨迹控制在保留非目标能力的同时实现了更持久的遗忘,并限制了推理开销。在三个基准和三个模型系列上,ST$^2$U实现了最强的整体平衡,结合了最佳或次佳的保留能力与有竞争力的遗忘效果,且受限知识重新进入的比例远低于测试时基线(13.76%-19.84% 对比 46.50%-59.10%)。

英文摘要

Controlling restricted knowledge in large language models is essential for model alignment and safe deployment. Test-time unlearning avoids costly retraining and parameter updates by intervening only during inference. However, existing activation-editing methods apply isolated pointwise corrections, overlooking how autoregressive generation continually reconstructs hidden states from the prompt, cache, and generated prefix. Consequently, later states may return to restricted knowledge regions after a locally successful correction, causing restricted knowledge re-entry. In this work, we propose Stateful Test-Time Unlearning via restricted knowledge boundary control (ST$^2$U), which formulates test-time unlearning as trajectory-wide boundary control. ST$^2$U first models restricted knowledge boundaries in low-dimensional invertible coordinates while leaving orthogonal non-target components unchanged. During inference, ST$^2$U monitors risk along the trajectory, applies minimal boundary corrections with contextual anchoring, and propagates historical correction states across tokens to mitigate knowledge re-entry. This trajectory-wide control enables more persistent forgetting while preserving non-target capabilities and limiting inference overhead. Across three benchmarks and three model families, ST$^2$U delivers the strongest overall balance, combining best or second-best retention with competitive forgetting and substantially less restricted-knowledge re-entry than test-time baselines (13.76%-19.84% versus 46.50%-59.10%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑