arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缩小缩放定律:资源受限大型语言模型中的参数效率与计算最优训练

Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models

Joe Dwyer

arXiv 2610.06387首次发表:更新:

发表机构

ECPI University(ECPI大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文综述了LLM缩放定律向计算最优训练的演变,聚焦参数与数据效率,提出在资源受限环境下应综合性能、参数、计算成本等多维度评估效率。

AI 中文摘要

大型语言模型(LLM)通过增加模型规模、训练数据和计算资源,已取得了显著的性能提升。然而,传统的缩放方法会产生收益递减、财务和环境成本上升,并为在大型工业实验室之外工作的研究人员设置了参与障碍。本综述考察了LLM缩放理论从经验缩放定律到计算最优训练的演变,特别强调了参数效率、令牌利用、数据效率和资源受限环境。综述综合了缩放定律的基础性工作,以及后续关于计算最优训练、数据剪枝、高效架构、量化、低秩适应和边缘导向优化的研究。文献表明,研究重点正从规模最大化转向对参数、令牌、计算和硬件资源的更审慎分配。同时,关于在企业级基础设施上建立的缩放原则是否能推广到较小模型和受限计算环境,仍存在重要的经验、理论和方法论空白。本综述将这些进展组织成一个资源高效LLM训练的统一框架,并认为未来的进展不应仅通过模型性能来评估效率,而应通过性能、参数数量、计算成本、令牌分配和硬件约束之间的关系来评估。

英文摘要

Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to compute-optimal training, with particular emphasis on parameter efficiency, token utilization, data efficiency, and resource-constrained environments. Foundational work on scaling laws is synthesized alongside later research on compute-optimal training, data pruning, efficient architectures, quantization, low-rank adaptation, and edge-oriented optimization. The literature indicates a shift from scale maximization toward more deliberate allocation of parameters, tokens, compute, and hardware resources. At the same time, important empirical, theoretical, and methodological gaps remain regarding whether scaling principles established on enterprise-grade infrastructure generalize to smaller models and constrained computing environments. This review organizes these developments into a unified framework for resource-efficient LLM training and argues that future progress should evaluate efficiency not solely through model performance, but through the relationship among performance, parameter count, computational cost, token allocation, and hardware constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑