arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08624cs.LGcs.AIstat.ML

平衡Adam的早期记忆选择

Early Memory Selection for Balanced Adam

Alberto Fernández-Hernández, Cristian Pérez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ortí

首次发表
浏览论文内容

中文总结 AI 辅助

提出从短试点训练中选取Adam共享记忆参数β的方法,利用局部模型平衡采样变异与延迟,估计三次方规则系数,在11个视觉和语言任务上显著降低验证差距。

中文摘要 AI 辅助

我们提出了一种方法,通过短暂的试点训练来选择Adam中共享的记忆参数$\eta_1=\eta_2=\eta$。选定的$\eta$在后续的完整训练中保持固定。Adam归一化方向的局部模型在采样变异性和由平均过去梯度引入的延迟之间取得平衡。这种平衡产生了一个三次方记忆规则,其两个系数通过几个试点检查点处的梯度探针来估计。该估计器联合使用分子和分母,保留它们的协方差。通过200次更新的试点,在四个检查点各使用十六个探针梯度,对十一个视觉和语言工作负载进行的种子匹配回顾性评估,相对于共享$\eta=0.95$的网格代表,平均相对验证差距降低了40.7%,最差四分之一的平均差距降低了44.3%。平均差距也比在所有十一个工作负载中选择的最佳常数$\eta$低32.3%。

英文摘要

We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training. The selected $β$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$. The mean gap is also 32.3% lower than that of the best constant $β$ chosen across all eleven workloads.

发表机构

  • Universitat Politècnica de València(瓦伦西亚理工大学)
  • Universitat Jaume I(海梅一世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑