arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12385cs.AI

双流式Transformer:将主预填路径与额外解码计算解耦

Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

Liming Liu, Mingze Wang, Tuo Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出双流式Transformer,将主预填路径与额外解码计算解耦,通过共享权重与缓存降低推理成本,在多架构与数据配置下实现更低验证损失,且可灵活分配MoE专家预算。

中文摘要 AI 辅助

随着大语言模型服务的请求增多,累积推理成本相较于一次性训练成本变得愈发重要。推理的两个阶段对硬件的压力各不相同:提示词预填是并行的,通常受计算限制;而自回归解码是顺序的,常受内存带宽限制。传统的宽度或深度缩放会同时增加两个阶段的成本,因为每添加一层都会在两个阶段中被计算。我们提出,能否将额外的学习计算分配给延续预测,同时保留提示词范围的主计算和单个持久键值(KV)缓存。我们提出了双流式Transformer。其主流是一个完整的因果语言模型,用于处理提示词并写入KV缓存;辅助流在提示词处理阶段被省略,仅从提示词的最终位置开始激活,添加延续预测计算,且不写入持久状态或影响主流。两个流共享大部分注意力、MLP和输出矩阵,同时使用独立的词嵌入和轻量耦合。权重共享和主缓存也为分组执行期间复用加载的权重及缓存的键值创造了机会。在匹配标记的对比中,双流式Transformer在不同架构和数据配置下均实现了更低的验证损失。在MoE模型中,这种分离使主专家和辅助专家的扇出成为对预填成本、延续成本和预测质量的独立控制。我们研究了两种场景:在固定预填专家计算量下增加解码计算,以及在两个流之间重新分配固定的解码专家预算。这些实验揭示了预填-解码质量的权衡,并证明了特定阶段专家分配的潜力。

英文摘要

As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at each decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving prompt-wide primary computation and a single KV cache. We realize this with the Decode-Branch Transformer. Its primary path alone processes the prompt and writes the KV cache; the decode branch is omitted during prefill and activated only from the final prompt position onward, adding continuation computation without writing state or affecting the primary path. The paths share attention, MLP, and output matrices, using separate token embeddings with lightweight coupling. Grouped decode reuses loaded weight tiles and the primary KV cache across both paths, so the added arithmetic does not proportionally increase dominant memory traffic or decode latency. Across matched-token comparisons, Decode-Branch achieves lower validation loss across architectures and data settings. In MoE models, the primary and branch expert fan-outs become independent knobs for trading prompt cost, decode cost, and predictive quality. We study two expert-allocation regimes, holding prefill or decode computation fixed, and expose a prefill-decode-quality trade-off enabled by phase-specific expert allocation.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑