发表机构
MBZUAI; National Taiwan University(穆罕默德·本·扎耶德人工智能大学; 国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TopK-Guided免训练方法,结合令牌级稀疏性适应与块级预算分配,在Llama模型上优于TEAL和WINA,提升推理效率。
AI 中文摘要
激活稀疏性通过将不重要的激活设为零来加速大型语言模型(LLM)推理,从而可以跳过相应的计算。然而,现有的免训练方法做出了不同的权衡:基于阈值的方法(如TEAL)使稀疏性水平适应每个令牌,但不能严格控制实际实现的稀疏性,而基于TopK的方法(如WINA)强制固定稀疏性水平,但对每个令牌使用相同的稀疏性预算。两者还在所有Transformer块上应用相同的预算,尽管块敏感性存在很大差异。我们引入了TopK-Guided,一种免训练方法,通过结合有界令牌级稀疏性适应和敏感性感知的块级预算分配来解决这两个局限性。在Llama-2和Llama-3模型上,TopK-Guided在困惑度和下游准确性方面始终优于TEAL和WINA,同时保持与WINA基本相同的稀疏性相关投影计算,在高稀疏性下收益最大。消融实验表明,这两个组件提供了互补的改进。
英文摘要
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.