arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPLASH: 面向LLM服务的注意力无缝切换并行布局

SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu, Qiuchu Yu, Xiyu Shi, Huimin Cui, Jiacheng Zhao

arXiv 2609.37626首次发表:更新:

发表机构

University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SPLASH提出一种在请求运行期间无缝切换注意力并行布局的服务系统,利用KV缓存与权重分片解耦实现近乎零开销的布局转换,显著提升LLM服务吞吐量。

AI 中文摘要

没有单一的注意力并行化方式能在所有负载下都很好地服务大型语言模型。低并发场景倾向于张量并行,大量独立请求倾向于数据并行注意力,而长提示词则倾向于上下文并行。推理、智能体和强化学习回滚等工作负载使得固定选择变得不可行:一个批次开始时包含许多短请求,结束时却变成少数极长请求,因此在同一批请求运行期间,最佳布局会发生变化。然而,服务引擎在启动时仍固定一种布局,因为更改布局意味着排空请求并重启工作节点。我们提出SPLASH,一个在请求运行期间切换注意力并行布局的服务系统。它基于一个观察:现代注意力机制,在KV头很少或没有的情况下,将请求的KV缓存存放位置与注意力权重的分片方式解耦。这带来两个后果。首先,不同布局仅在于谁拥有权重和缓存,且大部分状态已经位于下一个布局所需的位置;SPLASH重用这些状态,在后台进行中的推理中移动其余部分,并在批次边界处交接,使切换几乎免费:其中位开销低于其运行步骤的0.51%。其次,这种解耦暴露了一种现有引擎缺乏的布局:解耦所有权并行(DOP)像张量并行那样分片注意力权重,同时像数据并行注意力那样将每个请求的缓存保留在单一所有者上。DOP不复制任何一方,提供比数据并行注意力多27-60%的KV容量,并在KV内存限制准入时给调度器一个选择。一个感知转换的调度器在负载变化时跟随四种布局中的最佳者。在B200 GPU上服务GLM-5.3时,SPLASH相比固定布局部署将端到端服务吞吐量提升1.3-1.73倍,同样的布局机制也出现在H200上的DeepSeek-V3.2和DCU上的GLM-5.3-Flash中。

英文摘要

No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.

Comments28 pages, 11 figures, 8 tables. Code: https://github.com/ict-agent/SPLASH-sglang

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑