Transformer语言模型中的顺序计算蒸馏
Distilling Sequential Computation in Transformer Language Models
浏览论文内容
中文总结 AI 辅助
提出用轻量级合并模块将token序列折叠为替代嵌入,压缩Transformer输入与KV缓存,减少序列长度达40%且精度损失极小,适用于多种下游任务。
中文摘要 AI 辅助
Transformer语言模型以自回归方式逐token处理序列,这使得随着上下文增长,计算成本日益增加。然而,许多相邻的token片段高度可预测或经常作为稳定单元出现,表明它们的表示可能是可压缩的。我们提出了一种蒸馏顺序计算的方法,通过用轻量级合并模块动态计算的折叠表示替换输入token片段。该模块从一系列静态token嵌入中生成一个单一的替代嵌入,捕获多个token的功能角色,使预训练模型无需架构更改或重新训练即可处理压缩输入。我们在推理过程中应用此方法压缩提示和中间解码步骤,使用回滚机制将存储的多token KV缓存条目替换为其单步替代。跨多种模型的实验表明,合并模块可将有效序列长度减少高达40%,同时在语言建模评估和下游任务(包括问答、摘要、常识推理和长格式数学推理)中仅产生最小精度下降。对合并模块的额外轻量级适配进一步改善了选定设置中的精度-压缩权衡。这些结果表明,Transformer中的顺序token计算可以通过压缩的替代表示有效近似,这些表示无需模型更新即可逼近原始行为。
英文摘要
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module. This module generates a single surrogate embedding from a sequence of static token embeddings that captures the functional role of the multiple tokens, allowing pretrained models to operate on compressed inputs without architectural changes or re-training. We apply this approach during inference to compress both prompts and intermediate decoding steps, using a rollback mechanism to substitute stored multi-token KV cache entries with their single-step surrogates. Experiments across diverse models show that the merge module can be used to reduce effective sequence length by up to 40% with minimal accuracy degradation across language modeling evaluations and downstream tasks, including question answering, summarization, commonsense reasoning, and long-form mathematical reasoning. Additional lightweight adaptation of the merge module further improves the accuracy-compression trade-off in selected settings. These results demonstrate that sequential token computation in Transformers can be effectively approximated through condensed surrogate representations that approximate the original behavior without model updating.
发表机构
- The University of Chicago(芝加哥大学)
- Toyota Technological Institute at Chicago(芝加哥丰田技术学院)
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。