发表机构
Montclair State University(蒙特克莱尔州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数据许可证到期后预测模型中的记忆残留问题,本文系统基准测试了时间序列机器遗忘方法,发现2020年危机年份记忆差距最大,TSMixer配合hinge方法在多数场景下最接近重训练预言机,并强调了明确删除范围与功效感知审计的必要性。
AI 中文摘要
当数据许可证到期时,删除存储的记录并不能消除已训练预测器中编码的影响。机器遗忘旨在无需重训练的情况下消除这种影响。我们使用3,200个配对参考(基于全部数据训练)和预言机(在不包含指定期间的子集上重训练)对时间序列遗忘进行了基准测试。该网格覆盖了五种架构、四个滚动折叠、五个可删除年份,以及基于标普500波动率面板的三个实验性删除级别。2020年新冠疫情危机年份在所有架构中产生了最大的记忆差距。移除该年份在所有折叠中提升了所有三个可部署模型,其中在2022年熊市中提升最大,而两个不可部署模型的响应则不一致。近似遗忘的目标是预言机,而非在删除期间的低预测精度。在一个Transformer单元中,一个从未在2020年数据上训练的预言机仍能以0.51的信息系数预测该年份,而参考模型为0.55;将预测推向噪声会降低测试技能。在十二个可部署的架构-方法组合中,只有TSMixer配合hinge方法在每个折叠中仍接近预言机,弥合了参考与预言机之间74%-118%的差距,且没有可测量的测试技能损失。方法排名因架构和滚动窗口而异。审计分离度随先前的记忆程度而上升,但在精确删除后可能仍然很小。窗口级损失比较最多达到0.69,而将股票级窗口视为独立会使绝对t统计量中位数膨胀1.9倍。这些结果要求明确删除范围、针对相关架构和窗口的预言机验证,以及考虑功效的审计。
英文摘要
When a data license expires, deleting stored records does not remove influence encoded in a trained forecaster. Machine unlearning seeks to remove this influence without retraining. We benchmark temporal unlearning with 3,200 paired references trained on all data and oracles retrained without the requested period. The grid covers five architectures, four rolling folds, five deletable years, and three experimental deletion levels on an S&P 500 volatility panel. The 2020 COVID crisis year produces the largest memorization gap for every architecture. Removing it improves all three deployable models in every fold, with the largest improvement in the 2022 bear market, while the two non-deployable models respond inconsistently. The target for approximate unlearning is the oracle, not low predictive accuracy on the deleted period. In one Transformer cell, an oracle that never trained on 2020 still predicts it at an information coefficient of 0.51, compared with 0.55 for the reference; pushing predictions toward noise reduces test skill. Across twelve deployable architecture-method pairs, only TSMixer with the hinge method remains near the oracle in every fold, closing 74-118% of the reference-to-oracle gap without a measurable loss of test skill. Method rankings vary across architectures and rolling windows. Audit separation rises with prior memorization but can remain small after exact deletion. The window-level loss comparison reaches at most 0.69, and treating stock-level windows as independent inflates the absolute t-statistic by a median factor of 1.9. These results call for an explicit deletion scope, oracle validation for the relevant architecture and window, and power-aware auditing.