arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于矩估计的信赖域框架

A Trust-region Framework for Moment Estimation

Oluwasegun A. Somefun

arXiv 2608.04026首次发表:更新:

AI 中文总结

本文提出一种信赖域框架,推导得到\textsc{Gmake}机制,经GPT2-124M实验验证,不同强弱的信赖域约束下,二阶矩与四阶矩实现方式各有优势。

AI 中文摘要

本文提出一种信赖域框架,用于理解随机梯度优化中自适应矩估计机制(如\textsc{Adam})的行为。具体而言,该框架中每个权重的更新步长幅值被约束在由阶数$p\in[2,4]$的矩约束所控制的信赖域内。推导得到一系列基于二阶矩估计和归一化$p$阶矩估计的学习率机制,当$p=4$时涉及类似峰度的估计。该通用机制被称为\textsc{Gmake},在统一的信赖域框架下,为矩估计归一化、学习率调度、作为动量的谱低通滤波以及算子级谱归一化提供了统一解释。对在FineWeb-Edu和TinyStories上训练的GPT2-124M的实验表明,当信赖域约束较弱时,四阶矩实现方式的收益最大;随着引入的信赖域控制逐渐增强,二阶矩实现方式的竞争力不断提升,通常能达到比对应四阶矩实现方式略低的验证损失。

英文摘要

In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, the magnitude of the update step associated with each individual parameter is constrained by a finite-order $p$-moment trust-region, with $p\ge1$. The resulting derivation leads to a family of learning-rate mechanisms based on second-moment estimation and normalized $p$-th-moment estimation. For $p=4$, this involves kurtosis estimation. Subsequent derivations provide a unified interpretation of moment-estimation-based normalization, learning-rate scheduling, momentum as a spectral first-order lowpass regularization, and operator-level spectral-norm normalization within a common trust-region framework. Preliminary experiments on GPT2-124M trained on FineWeb-Edu and TinyStories suggest that the fourth-moment realization provides its greatest benefit when trust-region constraints are weak. As progressively stronger trust-region controls are introduced, the second-moment realization becomes increasingly competitive, often achieving slightly lower validation loss than its corresponding fourth-moment realization.

Comments20 pages, 5 figures. revised for improved presentation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑