arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29364cs.NI

NebulaSD:多对多推测解码

NebulaSD: Many-for-Many Speculative Decoding

Junhao He, Hongyang Du

首次发表
浏览论文内容

中文总结 AI 辅助

NebulaSD提出多对多推测解码系统,通过动态资源池和异步KV状态准备,在四GPU上提升请求处理率50.4%至72.6%,并提高GPU利用率。

中文摘要 AI 辅助

推测解码通过使用轻量级草稿模型提出候选令牌,供目标模型并行验证,从而加速大型语言模型(LLM)的推理。然而,草稿生成和验证表现出不同的服务特性,并偏好不同的批次配置,这使得在并发工作负载下固定的草稿-目标耦合效率低下。现有的分布式设计可以在物理上分离这两个阶段,但往往保留请求或批次亲和性,阻碍其容量在全球范围内共享。我们提出了NebulaSD,一种多对多(M-for-N)推测解码系统,它将草稿和目标工作器组织为独立可调度的资源池,并从共享请求池中动态重构阶段特定的批次。这种动态重新分配消除了固定的工作器局部性,要求请求状态在新选定的工作器上可用,而不引入迁移停顿。NebulaSD通过工作器触发的批次重构和与模型执行重叠的异步KV状态准备来应对这一挑战。我们从系统和扩展性两个角度评估了NebulaSD,结果表明,在四GPU部署中,动态池化相比物理分离的基线将请求轮次处理率提高了50.4%,相比同地执行提高了72.6%,同时大幅提高了有效GPU利用率。基于配置文件的模拟进一步表明,在理想化的状态移动下,计算侧容量扩展近似成比例。

英文摘要

Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Drafting and verification, however, exhibit different service characteristics and favor different batch configurations, making fixed draft-target coupling inefficient under concurrent workloads. Existing distributed designs can physically separate the two stages, but often retain request or batch affinities that prevent their capacities from being shared globally. We present NebulaSD, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools. Such dynamic reassignment removes fixed worker locality, requiring request states to be made available at newly selected workers without introducing migration stalls. NebulaSD addresses this challenge through worker-triggered batch reconstruction and asynchronous KV-state preparation overlapped with model execution. We evaluate NebulaSD from both system and scaling perspectives, showing that dynamic pooling improves request-round processing rate by 50.4% over a physically disaggregated baseline and 72.6% over co-located execution on a four-GPU deployment while substantially increasing effective GPU utilization. Profile-driven simulations further show approximately proportional compute-side capacity scaling under idealized state movement.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑