LLM压缩的能力缩减定律
Capability Scaling-Down Laws for LLM Compression
浏览论文内容
中文总结 AI 辅助
本研究系统探究LLM压缩(剪枝、量化、蒸馏)的能力缩减定律,提出简单预测关系,验证其准确性与泛化性,并展示其在压缩方法选择中的决策价值。
中文摘要 AI 辅助
LLM压缩降低了推理成本和内存需求,但选择方法和配置在很大程度上仍是经验性的,因为可比较的资源缩减可能产生不同的能力损失。我们系统地研究了跨剪枝、量化和蒸馏的LLM压缩的能力缩减定律。我们的框架衡量数学、代码生成和问答中的能力损失,并将这些测量与模型大小、训练阶段、压缩设置、数据可用性和训练暴露相关联。我们开发了简单的预测关系,并评估了它们的准确性、测量效率以及对未见配置和模型状态的泛化能力。跨剪枝级别共享密度响应,使拟合剪枝预测器所需的配置测量减半:在新的Pythia状态、预注册的OLMo-2测试状态以及Wanda剪枝下,紧凑关系与使用所有测量值在数学和代码上拟合的回归相匹配,每token差异在0.020 nats以内,系数针对每种设置重新拟合。受控蒸馏实验表明,重度数据复用的代价在问答分布中反复出现,而净收益取决于评估分布。我们进一步通过将数值选择与配置中位数和固定方法优先级进行比较来评估这些预测的决策价值。跨两个模型家族的独立评估表明,在测试的候选集内,选择捕获了问答中大部分可用的跨方法收益,其中固定方法优先级达到相同的遗憾,数学和代码的机会较小。这些结果阐明了能力缩减定律的预测范围及其在压缩方法选择中的用途。我们的代码公开可用,网址为:此https URL。
英文摘要
LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pruning, quantization, and distillation. Our framework measures capability loss in mathematics, code generation, and question answering, and relates these measurements to model size, training stage, compression settings, data availability, and training exposure. We develop simple predictive relations and evaluate their accuracy, measurement efficiency, and generalization to unseen configurations and model states. Sharing the density response across pruning levels halves the configuration measurements needed to fit a pruning predictor: on new Pythia states, on pre-registered OLMo-2 test states and under Wanda pruning, the compact relation matches a regression fitted with all measurements on math and code to within 0.020 nats per token, with coefficients refitted for each setting. Controlled distillation experiments show that the cost of heavy data reuse recurs across question-answering distributions, while the net benefit depends on the evaluation distribution. We further evaluate the decision value of these predictions by comparing numerical selection with configuration medians and fixed method priorities. Independent evaluations across two model families show that selection captures most of the available cross-method benefit for question answering within the tested candidate sets, where a fixed method priority attains the same regret, with smaller opportunities for mathematics and code. These results clarify the predictive scope of capability scaling-down laws and their use in compression method selection. Our code is publicly available at: https://github.com/LabRAI/scaling_down_law.
发表机构
- Florida State University(佛罗里达州立大学)
- Nokia Applied Research(诺基亚应用研究)
机构由 AI 辅助整理,请以论文原文为准。