轻量且通用的基于梯度历史重组的习得优化器
Lightweight and Versatile Learned Optimization by Recombination of Gradient History
浏览论文内容
中文总结 AI 辅助
本文提出一种轻量通用优化器,通过重组梯度历史(按时间跨度平均)减少预测空间,以37k参数网络在0.87 GPU小时内训练,零样本泛化并提升多种模型性能,FLOPs开销低至0.3%。
中文摘要 AI 辅助
本文提出了一种轻量且通用的习得优化器,它通过动态重组梯度历史(表示为不相交时间跨度上的平均值)来工作。该优化器将预测空间缩减为每个梯度平均值的一个标量系数,该系数由多个参数共享。对较旧梯度进行渐进平均可最小化长历史的内存成本,同时保持其贡献可独立访问。一个在0.87 GPU小时内训练完成的37k参数网络,能够零样本泛化到未见过的任务,在BERT-Tiny和GPT-Tiny上将验证损失分别降低了9.1%和0.4%,在Vision Transformer上将测试准确率较Adam提高了3.5个百分点,在九个图模型上平均提高了2.7个百分点,而FLOPs开销低至0.3%。
英文摘要
This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。