arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03432cs.LGcs.AI

OptiSelect:优化器如何塑造数据课程?

OptiSelect: How does the Optimizer Shape Data Curriculum?

Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出OptiSelect框架,理论证明优化器预条件器影响数据选择增益,实验表明AdamW评分几何最优,为LLM预训练中优化器与数据选择协同设计提供指导。

中文摘要 AI 辅助

在线数据选择通过在每个批次中训练最有价值的候选数据,已显著提升了大语言模型预训练的效率。由于候选数据的价值是通过其有效的模型更新来实现的,因此原则性的选择应考虑优化器步骤,该步骤在更新模型参数之前重塑原始梯度。我们将这种优化器感知的选择范式形式化为OptiSelect,并首次系统地研究了优化器如何塑造数据选择。我们的理论建立了一个选择增益原则,其中在线选择的优势由优化器诱导的效用分数的可区分性决定。我们证明了Lion和Muon的基于符号和极切向预条件器会遭受可区分性崩溃,从而限制了OptiSelect可获得的增益上限,而对角自适应优化器如AdamW和Sophia则具有严格更好的上界。所提出的原则还定量推导了最优候选数据过采样比率。在124M和720M模型上的预训练实验与我们的理论分析一致,并表明即使使用Muon作为优化器,AdamW的对角自适应评分几何仍是最强的评分几何。我们进一步证明,OptiSelect在数据改写(现代数据处理流程中使用的一种技术)下仍能保持其优势。我们的发现为在大语言模型预训练中共同设计优化器和数据选择提供了理论基础和实践指导。

英文摘要

Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate's value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first systematic study of how the optimizer shapes data selection. Our theory establishes a selection gain principle in which the advantage of online selection is governed by the discriminability of the optimizer-induced utility scores. We prove that sign-based and polar-tangential preconditioners of Lion and Muon would suffer from a discriminability collapse which caps attainable gains from OptiSelect, whereas diagonal-adaptive optimizers such as AdamW and Sophia admit strictly better upper bounds. The proposed principle also yields a quantitative derivation of the optimal candidate oversampling ratio. Pretraining experiments on 124M and 720M models are consistent with our theoretical analysis and show that AdamW's diagonal-adaptive scoring geometry remains the strongest scoring geometry even with Muon as optimizer. We further demonstrate that OptiSelect retains its benefits under data rephrasing, a technique used in modern data processing pipelines. Our findings provide theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.

发表机构

  • EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑