发表机构
University of Central Florida; Mohammed VI Polytechnic University(中佛罗里达大学; 穆罕默德六世理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TACO是一种新型优化器,通过逐列选择最大绝对值项的符号计算最速下降方向,大幅减少优化器状态内存,实现单GPU上30-32B模型的全参数微调,且保持精度和效率。
AI 中文摘要
大语言模型(LLM)的全参数微调会产生大量的优化器状态内存开销,限制了能在现代GPU上容纳的模型规模。现有方法要么压缩优化器状态,要么放弃一阶梯度,要么在保留稠密状态的同时改变更新几何结构。最近引入的Muon优化器通过矩阵值更新减少了优化器内存,但其几何结构与AdamW不同,在微调以AdamW预训练的模型时可能导致性能下降。为了在不牺牲LLM微调精度或计算效率的前提下减少优化器内存,我们提出了三元绝对最大值逐列单稀疏优化器(TACO),它遵循Muon的算子范数最速下降视角,但进一步推进了几何路径。TACO通过选择二维权重矩阵每列中绝对值最大项的符号,在维度归一化的$1\ o1$算子范数下计算精确的最速下降方向。这保留了一阶梯度,同时使优化器状态内存几乎可忽略。我们的实用TACO优化器每列仅维护少量低精度梯度分量,相对于AdamW8bit,持久优化器状态减少了$174\ imes$(从27.7 GB降至0.16 GB),在OPT-13B上峰值训练内存减少了$2.9\ imes$(从80.6 GB降至27.5 GB),同时实现了相当的精度和运行时间。TACO还支持在单个80 GB H100 GPU上对30-32B参数模型进行全参数微调,涵盖多个模型家族和任务。
英文摘要
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized $1\to1$ operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by $174\times$ relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by $2.9\times$ (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable accuracy and runtime. TACO further enables full-parameter fine-tuning of 30-32B-parameter models on a single 80 GB H100 GPU across multiple model families and tasks.
Comments24 pages, 7 figures, 10 tables. Code available at https://github.com/Jichao2357/TACO_optimizer