arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.00329cs.LGcs.AI

运行中Adam优化器的数据Shapley值

In-Run Data Shapley for Adam Optimizer

  • Mohamed bin Zayed University of Artificial Intelligence(莫萨德·本·泽德人工智能大学)
  • University at Buffalo(布法罗大学)
  • King Abdullah University of Science and Technology(卡布斯国王科学与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu

AI总结:

本研究提出Adam-aware运行中数据Shapley方法,通过线性化技术实现高精度数据归因,提升现代训练流程效率。

AI中文摘要:

可靠的数据归因对于减轻偏差和减少计算浪费在现代机器学习中至关重要,Shapley值作为理论黄金标准。虽然最近的

英文摘要:

Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard. While recent "In-Run" methods bypass the prohibitive cost of retraining by estimating contributions dynamically, they heavily rely on the linear structure of Stochastic Gradient Descent (SGD) and fail to capture the complex dynamics of adaptive optimizers like Adam. In this work, we demonstrate that data attribution is inherently optimizer-dependent: we show that SGD-based proxies diverge significantly from true contributions under Adam (Pearson $R \approx 0.11$), rendering them ineffective for modern training pipelines. To bridge this gap, we propose Adam-Aware In-Run Data Shapley. We derive a closed-form approximation that restores additivity by redefining utility under a fixed-state assumption and enable scalable computation via a novel Linearized Ghost Approximation. This technique linearizes the variance-dependent scaling term, allowing us to compute pairwise gradient dot-products without materializing per-sample gradients. Extensive experiments show that our method achieves near-perfect fidelity to ground-truth marginal contributions ($R > 0.99$) while retaining $\sim$95\% of standard training throughput. Furthermore, our Adam-aware attribution significantly outperforms SGD-based baselines in data attribution downstream tasks.

补充信息

↑