arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09444cs.LGcs.CLcs.DC

基于连续深度批处理的循环语言模型深度自适应推理

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对循环语言模型深度自适应推理中标准批处理失效的问题,提出连续深度批处理(CDB)方案,在两个模型上实现了接近理论上限的加速比,显著提升了吞吐量并降低了延迟。

中文摘要 AI 辅助

循环语言模型(LMs)的核心优势之一是深度自适应推理:通过对共享层模块进行可变次数的迭代,模型可对“简单”token使用更少计算资源,对“困难”token使用更多计算资源。但这种自适应特性打破了标准批处理机制:同一批次中的token现在需要不同次数的循环,因此不存在统一的前向传播,导致高效推理难以实现。vLLM等标准推理框架基于token级别调度,无法处理该问题,因为前向传播过程中需要从批次中移除token。已有研究提出了循环级调度作为解决方案,但始终未实现端到端方案。核心挑战在于循环架构还包含非循环的边界阶段(如token嵌入和LM头),其调度频率需与循环阶段不同。本文提出连续深度批处理(CDB),该方案以单个循环迭代为粒度进行调度,将边界阶段和循环步骤置于不同的优先级队列中,提前一步做出退出决策,并将所有调度工作与GPU计算重叠执行。在Ouro 1.4B和Huginn 3.5B模型上,CDB可实现自适应深度理论最大加速比的99%,转化为1.5至1.9倍的离线吞吐量提升,以及动态服务负载下45%至90%的归一化延迟降低。

英文摘要

A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps. Our method dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously. Experiments on Ouro 1.4B and Huginn 3.5B show that fully looped architectures are best suited to depth-adaptive inference, as large non-looped layers outside the recurrent core (e.g., token embedding, LM head, and unshared transformer blocks) slow down and complicate scheduling. Overall, CDB realizes up to 99% of the estimated maximum speedup available, leaving further gains primarily dependent on model architecture and exit behavior.

发表机构

  • Technical University of Munich (TUM)(慕尼黑工业大学(TUM))
  • Imperial College London(帝国理工学院)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
  • University of Potsdam(波茨坦大学)
  • Hasso Plattner Institute for Digital Engineering(哈索·普拉特纳数字工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑