发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MoARa通过模块感知秩分配与幅度-方向分解,在多种LLM架构上减少低秩预训练步数和墙钟时间,最高达41.7%步数减少。
AI 中文摘要
低秩梯度投影降低了大型语言模型(LLM)预训练中优化器状态的内存成本,但达到目标质量所需的步数和墙钟时间仍有显著的改进空间。我们将此归因于现有方法中的两个设计选择:投影秩预算在具有异质投影敏感性的Transformer模块间均匀分配,以及投影原始梯度会同时衰减其幅度和方向。我们提出MoARa,它将基于静态分析的模块感知投影秩分配与分块幅度-方向分解相结合;默认块大小设置在注意力头维度附近。在跨越Llama、Qwen和DeepSeek的五个Transformer架构(规模从300M到7B)上,带有MoARa的GaLore在Llama 2 7B上以少37%的步数和少34%的墙钟时间达到标准GaLore的最终困惑度,且在标准图编译下仅增加0.2%的峰值保留内存开销。在我们评估的六种低秩预训练方法中,仅模块感知秩分配就在所有六种方法上带来方向一致的步数减少。在兼容的主机上,双组件设计实现了高达41.7%的步数减少和37.1%的墙钟时间减少。
英文摘要
Low-rank gradient projection reduces the optimizer-state memory cost of large language model (LLM) pretraining, but the steps and wall-clock time needed to reach a target quality remain a meaningful axis for improvement. We attribute this to two design choices in existing methods: the projection-rank budget is allocated uniformly across Transformer modules with heterogeneous projection sensitivity, and projecting a raw gradient attenuates its magnitude and direction jointly. We propose MoARa, which combines a static profiling-based module-aware projection-rank allocation with a block-wise magnitude-direction decomposition; the default block size is set in the neighborhood of the attention head dimension. Across five Transformer architectures spanning Llama, Qwen, and DeepSeek at 300M to 7B scales, GaLore with MoARa reaches standard GaLore's final perplexity in 37% fewer steps and 34% less wall-clock time on Llama 2 7B, with only 0.2% peak reserved memory overhead under standard graph compilation. Across the six low-rank pretraining methods we evaluate, module-aware rank allocation alone delivers directionally consistent step reductions on all six. On compatible hosts, the two-component design reaches up to 41.7% step reduction and 37.1% wall-clock reduction.
CommentsAccepted to EMNLP 2026 Main Conference. 30 pages, 16 figures