arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Temperon:以少三分之一的挂钟时间实现全时SAM质量

Temperon: Full-Time SAM Quality at a Third Less Wall-Clock

Stamatis Mastromichalakis

arXiv 2609.17575首次发表:更新:

AI 中文总结

Temperon通过前43%周期用普通SGD探索、随后交接给SAM-Muon精炼器,在多个数据集上以少三分之一的挂钟时间达到全时SAM精度,并揭示了调度交接的原则性依据。

AI 中文摘要

锐度感知最小化(SAM)使每个训练步骤的成本翻倍,但其收益集中在训练结束时。我们研究了昂贵的训练模式应在何处使用,并提出了Temperon:一个在训练周期预算的前43%使用普通SGD探索器的方法,然后进行一次计划内的交接,将整个最终的余弦退火阶段交给一个包裹了SAM的Muon精炼器。在CIFAR-10/100、SVHN和Tiny ImageNet上(五个随机种子,时间报告为达到目标所需的周期数乘以空闲GPU校准的周期成本),Temperon在四个数据集中的三个上,在所有精度指标上与最佳的全时SAM方案持平,同时提前35%、34%和32%达到最难的共同目标,并且在同等成本下比已发表的SAM+SGD方案高出一个档次。消融研究使归因精确:在保持其他一切不变的情况下,Muon精炼器贡献了+0.85个百分点;探索器的形状及其重启没有贡献,我们将其从贡献中撤回。在匹配预算下重新运行最接近的竞争对手——后期SAM,显示了前沿:它在达到每个中等水平目标时最快,但Muon精炼器带来的档次提升(CIFAR-100上0.83,CIFAR-10上0.97)是任何SGD精炼方法在任何种子下都无法达到的,而在Tiny ImageNet上,Muon没有带来档次提升,竞争对手直接获胜——这是该方法的测量边界。这种分配规律迁移到GPT-2预训练(全时SAM质量,挂钟时间减少29%)和GLUE微调(在SAM成本的三分之一下,性能从不差于全时SAM)。两个常数组织了经济学:早期跳过SAM获得固定的信用,而一个Muon周期在所有四个数据集上的成本是SAM+SGD周期的1.50倍。最后,交接无法从轨迹中计时:在余弦调度下,精度曲线是先平台后激增,因此信息存在于调度中,这使得计划切换是有原则的而非方便的。代码和一个可通过pip安装的实现已发布。

英文摘要

Sharpness-aware minimization (SAM) doubles the cost of every training step, yet its benefit concentrates where training ends. We study where an expensive training mode should be spent and propose Temperon: a plain-SGD explorer for the first 43% of the epoch budget, then one scheduled hand-off that gives the entire final cosine anneal to a SAM-wrapped Muon refiner. On CIFAR-10/100, SVHN and Tiny ImageNet (five seeds, times reported as epochs-to-target times an idle-GPU-calibrated epoch cost), Temperon matches the best full-time-SAM recipe on accuracy everywhere while reaching the hardest common target 35%, 34% and 32% sooner on three of the four, and sits a tier above the published SAM+SGD recipe at level cost. Ablations make the attribution exact: the Muon refiner is worth +0.85pp with everything else fixed; the explorer's shape and its restarts are worth nothing, and we withdraw them as contributions. Re-running the closest rival, late-phase SAM, at matched budget shows the frontier: it is fastest to every mid-level target, but the tier the Muon refiner buys (0.83 on CIFAR-100, 0.97 on CIFAR-10) is reached by no SGD-refined method in any seed, and on Tiny ImageNet, where Muon buys no tier, the rival simply wins -- the measured boundary of the method. The allocation law transfers to GPT-2 pretraining (full-SAM quality at -29% wall-clock) and GLUE fine-tuning (never worse than full-time SAM at a third of its SAM cost). Two constants organize the economics: skipping SAM early buys a fixed credit, and a Muon epoch costs 1.50x a SAM+SGD epoch on all four datasets. Finally, the hand-off cannot be timed from the trajectory: under cosine schedules the accuracy curve is plateau-then-surge, so the information lives in the schedule, making the scheduled switch principled rather than convenient. Code and a pip-installable implementation are released.

Comments12 pages, 5 Figures, 4 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑