arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09266cs.LG

LionVote:针对Lion的逐层学习率自适应

LionVote: Per-Layer Learning Rate Adaptation for Lion

Kris Atallah

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对Lion在不同层参数有效规模差异问题,提出LionVote逐层学习率机制,通过特定诊断更新参数张量的复合水平,投票阈值源于多种因素,节奏有界且可选择。该机制在ViT-Tiny/CIFAR-100上提升了准确率,且适应值与架构、任务相关。

中文摘要 AI 辅助

逐层诊断显示,在规定学习率下,对于ViT-Tiny/CIFAR-100上的注意力和MLP参数,Lion的有效规模高2.6 - 2.8倍,对于归一化层高约2倍,这种32%的跨层类型差异无法通过单一全局速率再现。该测量来自LionVote,一种逐层学习率机制,每个参数张量维持一个复合水平,通过验证损失决胜器解决的两个诊断(梯度方向稳定性和动量健康)每c个epoch更新一个持久整数。投票阈值源于几何恒等式、指数移动平均时间常数和噪声底估计;节奏在结构上有界并通过消融选择。在ViT-Tiny/CIFAR-100上,LionVote达到69.7%的top-1准确率,优于Lion的69.0%(p < 0.02, Welch t检验)和AdamW的68.8%。逐层适应值取决于架构异质性和任务;在统一的CNN架构上,带余弦退火的调优SGD仍然占主导,在ViT架构上收益取决于任务。

英文摘要

Per-layer diagnostics reveal that, at the prescribed learning rate, Lion's effective scale is 2.6-2.8x too high for attention and MLP parameters and ~2x too high for normalization layers on ViT-Tiny/CIFAR-100; this 32% cross-layer-type disparity cannot be reproduced by a single global rate. The measurement comes from LionVote, a per-layer learning rate mechanism in which each parameter tensor maintains a compound level, a persistent integer updated every c epochs by two diagnostics (gradient direction stability and momentum health) resolved by a validation loss tiebreaker. Voting thresholds derive from geometric identities, the EMA time constant, and a noise-floor estimate; cadence is bounded structurally and selected by ablation. On ViT-Tiny/CIFAR-100, LionVote achieves 69.7% top-1 accuracy vs. Lion's 69.0% (p < 0.02, Welch's t-test) and AdamW's 68.8%. Per-layer adaptation value depends on both architectural heterogeneity and task; on uniform CNN architectures tuned SGD with cosine annealing remains dominant, and on ViT architectures gains are task-dependent.

发表机构

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑