显微镜下的梯度:对内存高效梯度计算方法的资源利用基准测试
Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods
浏览论文内容
中文总结 AI 辅助
该研究针对四种Transformer架构,测试五种梯度优化器在三种内存策略下的资源利用,发现梯度累积效果最优,Adam并非普遍最优,梯度检查点效果依赖架构,为模型训练部署提供实用指南。
中文摘要 AI 辅助
人工智能训练的资源强度不断上升,正给电力供应和碳预算带来压力,推动人们在受限硬件上开展内存高效训练的系统研究。我们在四种Transformer架构(ViT、ModernBERT、Llama 3.1 1B和NanoVLM)下,针对三种内存策略(标准训练、梯度检查点和梯度累积),对五种梯度优化器(SGD、Adam、Adagrad、Adadelta和共轭梯度下降)进行基准测试,测量训练损失、GPU利用率、训练时间和内存使用情况。梯度累积是最可靠的策略,它在视觉-语言模型上将训练损失降低约一个数量级,在语言模型上降低约四倍,且无需额外GPU内存。与常见做法相反,Adam并非普遍最优:Adadelta和SGD在编码器和自回归架构上的表现优于它。梯度检查点的效果高度依赖架构:它能降低视觉Transformer的损失,但会严重降低编码器模型的性能,且在内存受限模型上会使训练时间增加最多60%。GPU利用率主要由架构决定,内存受限的语言模型的GPU利用率为8%-15%,计算受限的视觉模型则为96%-99%。这些发现为资源高效型模型的训练和部署中优化器与梯度策略的选择提供了实用指南。
英文摘要
AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware. We benchmark five gradient optimizers (SGD, Adam, Adagrad, Adadelta, and Conjugate Gradient Descent) under three memory strategies (standard training, gradient checkpointing, and gradient accumulation) across four transformer architectures (ViT, ModernBERT, Llama 3.1 1B, and NanoVLM), measuring training loss, GPU utilization, training time, and memory usage. Gradient accumulation emerges as the most reliable strategy, cutting training loss by roughly an order of magnitude on the vision-language model and about four-fold on the language model without additional GPU memory. Contrary to common practice, Adam is not universally superior: Adadelta and SGD outperform it on the encoder and autoregressive architectures. Gradient checkpointing's effect is strongly architecture-dependent, improving vision transformer loss while severely degrading the encoder model, and it increases training time by up to 60% on memory-bound models. GPU utilization is governed primarily by architecture, ranging from 8-15% for the memory-bound language model to 96-99% for compute-bound vision models. These findings provide practical guidelines for optimizer and gradient-strategy selection in resource-efficient model training and deployment.