arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

内动量用于差分隐私Muon

Inner Momentum for Differentially Private Muon

Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai

arXiv 2610.02738首次发表:更新:

发表机构

Los Alamos National Laboratory; University of Missouri; Georgia Institute of Technology(洛斯阿拉莫斯国家实验室; 密苏里大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对差分隐私训练中裁剪扭曲Muon更新几何的问题,提出在裁剪前对样本梯度进行历史平均,以限制裁剪失真并保持极因子,在私有GPT-2微调中提升BLEU和ROUGE-L。

AI 中文摘要

差分隐私训练在添加噪声前会对每个样本的梯度进行裁剪。这种裁剪对每个样本是径向的,但不等的裁剪因子会扭曲其平均值的相对奇异向量几何结构。Muon特别容易受到此影响,因为其更新是近似极因子UV^T,该因子仅依赖于裁剪可能改变的奇异向量。为了抑制这种退化,我们提出在裁剪前,对每个采样样本的Muon梯度在当前模型和近期模型的短历史上进行平均。裁剪后的批量矩阵随后分解为公共重缩放和采样梯度与裁剪值之间的协方差残差R,其中||R||_F <= sigma_lambda sigma_G,直接限制了裁剪引起的失真。我们进一步证明,在这些谱条件下,有限次Newton-Schulz迭代保持其输入的极因子,确认我们的修正能在正交化后幸存。在E2E和DART上对私有GPT-2进行微调,epsilon在{1, 2, 4, 8}时,DP-Muon-IM在每次种子匹配比较中均优于DP-Muon的BLEU和ROUGE-L,非私有诊断显示预噪声极误差降低2-4%。

英文摘要

Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.

CommentsLA-UR Number: LA-UR-26-28799

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑