arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13922cs.LGcs.DC

小批量持久化,八年之后:批量重用以步数和焦耳计的成本,以及它在数据上的节省

Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data

Matteo Fischetti

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过预注册实验系统评估小批量持久化技术,发现其仅在数据稀缺且大批量时节省数据,不节省步数、时间或能量,并建议在数据可重读时采用间隔轮次。

中文摘要 AI 辅助

小批量持久化重用数据而非读取数据:它不是在每个优化器步骤中抽取新的小批量,而是在同一小批量上连续执行K步。该方法在2019年被数据回显所吸收,但一直存在一个反对意见——重用只是模仿更大的学习率——且没有基线像方法本身那样经过仔细调优。本文进行了缺失的测试。一项预注册研究在FineWeb-Edu上训练了一个4900万参数的Transformer,小批量大小B取自{32, 128, 512},每个单元8个随机种子,针对每个批量大小和每个实验臂分别调整学习率,并与一个不改变采样且无重用的对照组进行比较。每个主要结论都是达到固定损失的成本,从四个维度衡量:优化器步数、新令牌数、秒数和插座处的焦耳数。然后我们在新随机种子和更新的GPU代上重复实验,并将三个实验臂置于统一的步数调度上。我们的研究结果是,小批量重用所获得的既不是速度也不是能量,而是数据,且仅在大批量大小下如此:在B=32时,它读取的新令牌比基线更多,而非更少。在步数、秒数和焦耳数上,它至多是免费的;而在B=512时,它表现最佳,但一个注册对照在n=8个随机种子下无法将重用的效果与学习率调度中的位置区分开来。因此,该技术值得在新鲜数据而非计算是主要瓶颈的情况下使用:如数据耗尽的数据集、按样本付费的流水线或无法回退的数据流。在数据可以简单重新读取的情况下,间隔轮次的效果与之相当或更好。

英文摘要

Minibatch persistency reuses data instead of reading it: rather than drawing a fresh minibatch at every optimizer step, it takes K consecutive steps on the same one. Absorbed into data echoing in 2019, it has carried one objection -- that reuse merely imitates a larger learning rate -- and no baseline tuned as carefully as the method itself. This paper runs the missing test. A pre-registered study trains a 49M-parameter Transformer on FineWeb-Edu at minibatch size B in {32, 128, 512}, 8 seeds per cell, tuning the learning rate separately for every batch size and every arm, against a reuse-free control that changes the sampling and nothing else. Each headline claim is a cost to reach a fixed loss, read on four axes: optimizer steps, fresh tokens, seconds, and joules at the socket. We then replicate on new seeds and a newer GPU generation, and put the three arms on one schedule in steps. The outcome of our study is that what minibatch reuse buys is neither speed nor energy but data, and only at large minibatch size: at B = 32 it reads more fresh tokens than the baseline, not fewer. On steps, seconds and joules it is at best free; and at B = 512, where it looks best, a registered control cannot separate the effect of reuse from the position on the learning-rate schedule at n = 8 seeds. The technique is therefore worth using where fresh data rather than compute is the binding cost: a corpus that runs out, a pipeline that pays per sample, a stream that cannot be rewound. Where the data can simply be read again, spaced epochs do as well or better.

发表机构

  • University of Padova(帕多瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑