LRCC:基于条件计算泛化低秩压缩
LRCC: Generalizing Low-Rank Compression with Conditional Computation
浏览论文内容
中文总结 AI 辅助
LRCC通过训练轻量级路由器为每个令牌选择嵌套低秩路径,在冻结低秩因子下优化路由器,在相同预算下提升语言模型性能,如Llama-2-7B下游准确率提高7.6个百分点。
中文摘要 AI 辅助
低秩压缩通过将线性变换替换为低秩分解来降低预训练语言模型的成本。然而,传统方法在推理期间使用固定的秩分配,无论输入令牌如何,都分配相同的计算量。我们引入了低秩条件计算(LRCC),通过为每个Transformer块训练一个轻量级路由器,从一小套嵌套低秩路径中选择,从而向预训练模型添加令牌相关的计算。在训练期间,低秩因子保持冻结,仅优化路由器。我们在Llama和Qwen模型上评估了LRCC,用于语言建模和零样本下游任务。在相同的平均活跃参数预算内,LRCC相比静态低秩压缩提高了预测性能,包括在Llama-2-7B上平均下游准确率比静态方法提高了7.6个百分点。在匹配的批大小-1解码延迟下,LRCC在Llama-3.2-1B上同时提高了困惑度和下游准确率,并在Llama-2-7B上保持竞争力,且无需专门的内核。最后,我们通过分析路由器的路径选择来评估分配令牌级路径的实用性。
英文摘要
Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.
发表机构
- Sapienza University of Rome(罗马第一大学)
机构由 AI 辅助整理,请以论文原文为准。