arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29711cs.LGcs.AI

解耦知识与隐私:面向大语言模型持续学习的任务后自蒸馏回放

Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning

发表机构南京航空航天大学 · 南京大学 · 清华大学
查看机构详情
  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
  • Nanjing University(南京大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Shengtao Wen, Yunying Yang, Xiang Chen, Lingbing Guo, Yu Tian, Sheng-Jun Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对隐私保护持续学习中保留与纠正的粒度冲突,提出SPARK分解方法,先冻结任务后分布再选择性纠正,实现有效PII抑制并保持知识保留。

中文摘要 AI 辅助

隐私保护的持续学习(PPCL)必须在保留跨顺序任务有用知识的同时,减少敏感内容的再现。形式化的隐私保证刻画了随机化机制,而操作性输出控制则关注训练后的模型是否选择性地降低其输出中敏感内容的可能性。在本工作中,我们在现实任务演化背景下,将后者与持续学习效用一同研究。保留与隐私纠正在不同粒度上运作:任务获取需要对当前和旧任务行为进行广泛保留,而隐私纠正则针对稀疏的标注位置。联合优化使得当前任务的保留目标持续变化。我们提出SPARK,一种保留-纠正分解方法,首先冻结学习到的任务后分布,然后在此稳定参考点周围应用选择性纠正。自蒸馏回放学习当前任务,同时蒸馏先前任务的行为,而任务后隐私纠正在锚定当前和旧任务非PII行为到最终检查点的同时,降低标注PII的可能性。广泛评估表明,SPARK在不同设置下实现了有效的选择性PII抑制,同时保持了强大的持续学习效用和知识保留。代码和数据将在发表后发布。

英文摘要

Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output control concerns whether a trained model selectively reduces the likelihood of sensitive content in its outputs. In this work, we investigate the latter together with continual-learning utility under realistic task evolution. Retention and privacy correction operate at different granularities: task acquisition requires broad preservation of current- and old-task behavior, whereas privacy correction targets sparse annotated positions. Joint optimization leaves the current-task preservation target continually changing. We propose SPARK, a retention-correction decomposition that first freezes the learned post-task distribution and then applies selective correction around this stable reference. Self-Distillation Replay learns the current task while distilling behavior from previous tasks, and Post-Task Privacy Correction reduces annotated-PII likelihood while anchoring current- and old-task non-PII behavior to the resulting checkpoint. Extensive evaluations demonstrate that SPARK achieves effective selective PII suppression while preserving strong continual-learning utility and knowledge retention across diverse settings. Code and data will be released upon publication.

↑