发表机构
The Chinese University of Hong Kong, Shenzhen; Zhejiang University; University of Sydney; Northwestern University; Analogy AI, Inc.; Mohamed Bin Zayed University of Artificial Intelligence(香港中文大学(深圳); 浙江大学; 悉尼大学; 西北大学; Analogy AI公司; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对时间序列预测中机器遗忘的挑战,提出RDTU残差扩散框架,通过神经正切核基础预测与扩散残差校正生成伪标签,实现轻量级模型更新,实验证明其效果最接近精确重训练。
AI 中文摘要
时间序列预测广泛用于敏感领域。这些场景中的模型通常在纵向的用户级或实体级记录上进行训练,这些记录可能随后因包含敏感或专有信息,或因传感器故障而损坏,需要被移除。为了在不进行昂贵重训练的情况下处理此类删除请求,机器遗忘作为一种隐私保护和数据治理的实用机制已被广泛研究。然而,机器遗忘在时间序列预测中的应用尚未得到很好的实现;这主要归因于以下独特挑战:基于梯度的遗忘可能不稳定,因为一个被删除的观测参与多个因果连接的预测窗口,导致参数更新传播超出请求的时间区间,并降低保留的预测效用。标签引导的更新提供了一种更可控的替代方案,但连续且上下文相关的预测缺乏合适的替换目标,而精确重训练的输出在遗忘过程中不可用。此外,被删除的时间模式在剩余数据中的支持高度不均匀。一些受影响的窗口在保留数据中存在结构相似的对应部分,而其他窗口则变得代表性不足或孤立。我们提出了RDTU,一个用于时间序列遗忘的残差扩散框架。RDTU首先使用保留集的神经正切核预测器获得一个与删除兼容的基础预测。然后,它利用保留参考数据的体积贡献来量化每个受影响窗口的全局和局部结构支持。随后,一个扩散模型生成残差校正,估计反事实预测,产生一个伪标签场,指导轻量级模型更新。实验表明,RDTU始终产生与精确重训练最接近的遗忘模型。
英文摘要
Time-series forecasting is widely used in sensitive domains. Models in these settings are often trained on longitudinal user- or entity-level records, which may later require removal because they contain sensitive or proprietary information or have been corrupted by sensor failures. To address such deletion requests without costly retraining, machine unlearning has been widely studied as a practical mechanism for privacy protection and data governance. However, the application of machine unlearning to time series prediction has not yet been well realized; this is mainly due to the following unique challenges: Gradient-based unlearning can be unstable because a deleted observation participates in multiple causally connected forecasting windows, causing parameter updates to propagate beyond the requested interval and degrade retained forecasting utility. Label-guided updating offers a more controlled alternative, but continuous and context-dependent forecasts lack a suitable replacement target, while the exact-retrained output is unavailable during unlearning. Moreover, the remaining support for a deleted temporal pattern is highly non-uniform. Some affected windows retain structurally similar counterparts in the retained data, whereas others become underrepresented or isolated. We present RDTU, a Residual Diffusion framework for time-series unlearning. RDTU first uses a retained-set neural tangent kernel predictor to obtain a deletion-compatible base forecast. Then it quantifies the global and local structural support of each affected window using the volume contribution of the retained-reference data. Then a diffusion model generates a residual correction that estimates the counterfactual forecast, yielding a pseudo-label field that guides a lightweight model update. Experiments show that RDTU consistently produces unlearned models that most closely match exact retraining.
Comments22 pages