AI 中文总结
针对内存有限设备上的自回归小型语言模型,提出结合信息保留与吞吐量增益的复合层重要性度量,可调整速度-质量权衡,预测加速比误差约4%,性能优于部分同类方法。
AI 中文摘要
小型语言模型(sLLMs)如今部署在内存和计算预算有限的设备上。在自回归设置中,推理受内存带宽限制:均匀量化往往对这类模型有害,因为其架构冗余度有限,且仅有少数层对低精度不敏感。我们提出一种复合度量,结合两个正交准则:信息保留(以归一化SQNR系数衡量)和吞吐量增益(用基于Roofline的延迟分析建模)。通过对Gemma 3 1B进行分析,我们发现前馈网络模块和嵌入矩阵是最具潜力的加速目标。对每个候选目标,我们基于模拟量化估计归一化质量分数,基于Roofline建模估计归一化速度分数,无需实际执行。我们将两个分数结合为复合优先级系数,可按需调整速度与质量的权衡。该度量具有通用性,可用于对单个模块、其投影子层或整个Transformer层进行优先级排序。我们在多种模型架构上评估了该方法,结果表明,我们的估计对加速比的预测误差约为4%。与需要昂贵近似推理的进化搜索、专用加速器或基于Shapley值的方法相比,我们的方法通常会为最具表达力的层分配更多资源。我们的分析方法使sLLM量化成为一项可预测的工程任务。
英文摘要
Small language models (sLLMs) are nowadays hosted on devices with limited memory and computational budget. In an autoregressive setup, inference is memory-bandwidth bound: uniform quantization is often detrimental to such models, since their architecture has limited redundancies and only a few layers are not very sensitive to lower precision. We propose a composite metric that combines two orthogonal criteria: information retention (measured in terms of a normalized SQNR-based coefficient) and throughput gains (modeled using a roofline-based latency analysis). By profiling Gemma 3 1B, we find that Feed-Forward Network blocks and the embedding matrix are the most promising targets for acceleration. For each candidate, we estimate a normalized quality score based on simulated quantization and a normalized speed score based on roofline modeling with no actual execution needed. We combine the two scores in a composite priority coefficient, allowing us to tune the trade-off between speed and quality as needed. Our metric is general and can be used to prioritize individual blocks, their projection sublayers, or transformer layers as a whole. We evaluate our approach on several model architectures, showing that our estimates have at around 4% prediction error for the accelerated speedup. We find that our method generally allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference. Our analytical approach makes sLLM quantization a predictable engineering task.
Comments25 pages, 11 figures