arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14107cs.CLcs.AI

北极星:用于扩散语言模型高效推理的漂移感知缓存校准和令牌承诺

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

Mingyu Lee, Akshat Ramachandran, Souvik Kundu, Tushar Krishna

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对扩散语言模型推理效率受双向注意力和静态阈值影响的问题,提出北极星框架,利用令牌表示漂移信号,由北极星缓存和北极星提交两组件构成,大幅提升了模型在准确性-吞吐量方面的表现及解码并行性。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)的推理效率受到两个挑战的限制:双向注意力妨碍了高效的KV缓存重用,而使用静态置信阈值增加解码并行性可能会损害生成质量。我们发现这两个挑战都源于一个共同现象:随着令牌被解码,通过双向注意力的上下文整合会导致令牌表示在解码步骤中漂移(演变)。基于此,我们提出了北极星,一个无需训练的推理框架,它使用令牌表示漂移作为统一信号来共同应对这两个挑战。北极星由两个组件组成:北极星缓存,通过漂移识别过时的KV缓存位置并执行稀疏的KV缓存刷新以实现高效重用;北极星提交,检测急剧漂移事件以可靠地识别准备提交的令牌。在几个dLLM系列的数学和编码基准测试中,北极星在准确性-吞吐量帕累托前沿上创造了新的技术水平,与现有基线相比,准确性提高了10.73%,吞吐量提高了3.7倍,并且在前向传递中实现了3.67个令牌的高解码并行性。

英文摘要

The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality. We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token representations to drift (evolve) across decoding steps. This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges. Polestar comprises two components: Polestar-Cache, which identifies stale KV-cache positions via drift and performs sparse KV-cache refreshes to enable efficient reuse, and Polestar-Commit, which detects sharp drift events to reliably identify commit-ready tokens. Across mathematics and coding benchmarks on several dLLM families, Polestar sets a new state of the art on the accuracy-throughput Pareto frontier, achieving up to 10.73% accuracy improvement, up to 3.7x higher throughput, and high decoding parallelism of 3.67 tokens per forward pass over existing baselines.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • Intel AI Group(英特尔人工智能集团)

机构由 AI 辅助整理,请以论文原文为准。

↑