arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10630cs.LGcs.AI

Phase-HDC:在离散相位学习中用梯度阈值替代优化器历史记录

Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning

Ahmed Nebli

首次发表
浏览论文内容

中文总结 AI 辅助

Phase-HDC用梯度阈值替代优化器历史记录,在离散相位学习中大幅降低存储需求,精度优于8比特Adam且接近6比特Adam,在多类数据集上表现突出。

中文摘要 AI 辅助

训练紧凑模型所需的内存往往远多于存储模型本身,因为优化器会保留自身的过去梯度记录。对于一种学习参数为低比特角度的超维分类器(我们称之为相位记忆),这些记录占用的内存是模型本身的数倍。我们探究是否能在仅存储模型的情况下训练此类模型。所提方法Phase-HDC在每次更新时,仅当当前梯度足够大时,才会按梯度符号将每个存储的角度调整至多一步。我们证明,这一简单规则是一阶损失模型的精确解,其中每个被更改的参数需支付固定成本。在除更新规则外其余条件均固定时,Phase-HDC的精度与使用6比特矩的Adam相当,但存储量仅为其三分之一。在11个图像、表格和文本数据集上,它的存储量比标准float32 Adam少16至23倍,比8比特Adam少4至6倍。代价是与float32 Adam相比,平均精度损失约5个百分点;而在11个数据集中的6个(包括字节级文本预测,其中8比特Adam会失效),Phase-HDC的精度高于8比特Adam。通过带检测的训练运行可解释这些结果:当参数必须位于离散网格上时,Adam的矩主要决定参数是否移动,这一决策可通过无需内存的当前梯度阈值实现,而矩的粗量化会破坏对罕见输入的决策。存储节省的是逻辑状态而非实测硬件内存。

英文摘要

Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients. For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a \emph{phase memory}, these records take several times more memory than the model itself. We ask whether such a model can be trained while storing nothing but the model. The proposed method, Phase-HDC, turns each stored angle by at most one step per update, against the sign of its current gradient, and only when that gradient is large enough. We show that this simple rule is the exact solution of a first-order loss model in which every changed parameter pays a fixed cost. When everything except the update rule is held fixed, Phase-HDC matches the accuracy of Adam with 6-bit moments while storing three times less. Across eleven image, tabular, and text datasets, it stores 16--23$\times$ less than standard float32 Adam and 4--6$\times$ less than 8-bit Adam. The price is an average loss of about five accuracy points against float32 Adam, while Phase-HDC is more accurate than 8-bit Adam on six of the eleven datasets, including byte-level text prediction, where 8-bit Adam collapses. Instrumented training runs explain these outcomes. Once parameters must sit on a discrete grid, Adam's moments mainly decide whether a parameter moves at all, a decision that a threshold on the current gradient can make without memory, and coarse quantization of the moments breaks this decision for inputs that the data rarely contain. The storage savings are logical state rather than measured hardware memory.

发表机构

  • Mathalyse Research(马萨莱研究机构)

机构由 AI 辅助整理,请以论文原文为准。

↑