arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Clean:通过Nyström Sketching实现线性内存成本的二阶LLM训练

Clean: Second-order LLM Training at Linear Memory Cost via Nyström Sketching

Beheshteh T. Rakhshan, Sahar Rajabi, Maziar Sargordi Shikai Fang, Guillaume Rabusseau, Sirisha Rambhatla

arXiv 2610.04204首次发表:更新:

发表机构

Mila & DIRO, Université de Montréal; Critical ML, University of Waterloo; Zhejiang University; Canada CIFAR AI Chair(蒙特利尔大学米拉与DIRO; 滑铁卢大学Critical ML; 浙江大学; 加拿大CIFAR AI主席)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Clean提出一种基于Nyström Sketching的全曲率优化器,将内存复杂度从二次降至线性,并支持单GPU训练13B模型,比AdamW快26%且内存更小。

AI 中文摘要

训练大型语言模型(LLM)面临一个基本权衡:内存高效的优化器(如Adam)丢弃了跨参数曲率信息,而全曲率方法(如SOAP)虽然能加速收敛,但内存成本高得令人望而却步。我们提出Clean,一种内存高效且全曲率的优化器,旨在解决这一瓶颈。Clean利用随机化Nyström方法精确近似SOAP中的左、右预处理器,将优化器的内存复杂度从模型维度的二次方降低到线性。随后,我们重新整合子空间外的分量,以捕获低秩近似之外的曲率信息,在最小内存成本下保留丰富的曲率。我们进一步提出Q-Clean,一种低精度变体,积极压缩优化器状态。在预训练LLaMA-1.3B架构时,与Muon相比,Q-Clean将优化器内存消耗减少了超过50%,同时保持强大且具有竞争力的预测性能。值得注意的是,Clean在优化器状态占用上比标准AdamW更小,同时达到AdamW的最终性能在墙钟时间上快26%。此外,我们的方法独特地支持在单个80GB GPU上预训练130亿参数的模型,为大规模模型优化提供了一种可扩展、高效且易用的方法。

英文摘要

Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right preconditioners in SOAP, and to reduce the optimizer's memory complexity from quadratic to linear in terms of model dimensions. We subsequently reintegrate the off-subspace components to capture curvature information beyond the low-rank approximation, preserving rich curvature at minimal memory cost. We further propose Q-Clean, a low-precision variant that aggressively compresses optimizer states. Q-Clean reduces optimizer memory consumption by \textbf{over 50\%} compared to Muon when pre-training a LLaMA-1.3B architecture, all while maintaining strong and competitive predictive performance. Notably, Clean operates with a smaller optimizer-state footprint than standard AdamW while reaching AdamW's final performance \textbf{26\% faster} in wall-clock time. Furthermore, our methods uniquely enable the pre-training of a 13B-parameter model on a single 80GB GPU, providing a scalable, efficient, and accessible approach to large-scale model optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑