不要丢弃Dropout:优化层稀疏性以实现高效的大语言模型训练与推理
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
浏览论文内容
中文总结 AI 辅助
本研究针对LLM预训练中缺失层Dropout的问题,通过超2400项实验,优化层Dropout的关键参数,证明其可降低训练FLOPs、提升效率并实现1.5倍推理加速,为高效LLM训练与推理提供了最佳实践。
中文摘要 AI 辅助
层Dropout(又称随机深度)已被证明可在语言和视觉Transformer中实现更快的训练、更高的准确率以及对零样本层剪枝的鲁棒性。然而,随着模型和数据集规模的扩大,Dropout——尤其是层Dropout——在大语言模型(LLM)的预训练方案中已基本消失。尽管一些前期工作报告称Dropout会降低准确率,但尚无全面研究量化该影响,更未对其进行缓解。本研究表明,层Dropout应应用于最先进的LLM训练中,为训练及训练后收益建立最佳实践与缩放分析。具体而言,通过优化层分布、时间调度和优化器超参数,我们发现在相同训练FLOPs下,层Dropout可降低损失;在给定训练步数下,LLM可实现更低或相近的验证损失,同时节省高达25%的训练FLOPs。此外,层Dropout支持显著的训练后优化,如早退出、中间层跳过和自投机解码,可实现高达1.5倍的推理加速,且准确率损失可忽略不计。我们开展了超过2400项训练实验,涉及参数规模从2.71亿到82亿的模型及多达1600亿token的数据集,证明这些发现可可靠扩展至大规模训练场景。所有预训练实验均在Cerebras CS-3系统上运行。
英文摘要
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same training FLOPs layer dropout leads to lower loss. For a given number of training steps, LLMs can achieve lower or similar validation loss while saving upto 25% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.5x inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.
发表机构
- Cerebras Systems(瑟布雷布拉斯系统公司)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。