发表机构
The University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
vSkipper通过虚拟化层将动态层跳过集成到现代服务引擎,保留连续批处理等特性,在负载拐点降低延迟36.8%(GSM8K),饱和时吞吐提升11.3%,首个实现每token内部层跳过服务收益的系统。
AI 中文摘要
动态层跳过通过允许每个token仅执行模型层的一个子集来减少LLM计算量。然而,现有的跳过器依赖于专门的生成循环,并未与现代服务引擎集成。因此,执行的层数减少并不一定能转化为更低的服务延迟:FlexiDepth平均跳过Llama-3-8B的32层中的8层,但其标准生成循环的解码速度比基础模型慢14.6%至21.0%。我们提出了vSkipper,一个虚拟化层,使得动态层跳过器能够在服务引擎中即插即用,同时保留连续批处理、固定形状批次、分页KV缓存和捕获的解码图。在每个路由层,vSkipper根据跳过器的决策对token进行分组,并仅在预测有利可图时使用路由执行。我们在SGLang中实现了vSkipper,并在相同提示、到达时间、输出长度和启动设置下,将发布的FlexiDepth检查点与上游SGLang进行评估。在上游负载曲线的拐点处,vSkipper在GSM8K上将平均端到端延迟降低了36.8%,在BBH上降低了13.6%。在饱和状态下,它将请求吞吐量提高了11.3%和7.4%。服务过程除了检查点自身的质量损失外,没有引入统计上可解析的质量损失。跨合成跳过策略、两个Qwen3跳过器和三个GPU,我们展示了无需工作负载特定调优的复用性。据我们所知,vSkipper是第一个在现代LLM服务引擎中实现每个token内部层跳过的服务效率收益的系统。代码作为SGLang的分支在此https URL开源。
英文摘要
Dynamic layer skipping reduces LLM computation by allowing each token to execute only a subset of the model's layers. However, existing skippers rely on specialized generation loops and do not integrate with modern serving engines. As a result, fewer executed layers do not necessarily translate into lower serving latency: FlexiDepth skips 8 of Llama-3-8B's 32 layers on average, yet its standard generation loop decodes 14.6--21.0% more slowly than the base model. We present vSkipper, a virtualization layer that makes dynamic layer skippers pluggable in serving engines while preserving continuous batching, fixed-shape batches, paged KV caching, and captured decode graphs. At each routed layer, vSkipper groups tokens by the skipper's decision and uses routed execution only when predicted to be profitable. We implement vSkipper in SGLang and evaluate the released FlexiDepth checkpoint against upstream SGLang under identical prompts, arrivals, output lengths, and launch settings. At the knee of upstream's load curve, vSkipper reduces mean end-to-end latency by 36.8% on GSM8K and 13.6% on BBH. Under saturation, it increases request throughput by 11.3% and 7.4%. Serving adds no statistically resolved quality loss beyond the checkpoint's own. Across synthetic skip policies, two Qwen3 skippers, and three GPUs, we demonstrate reuse without workload-specific tuning. To our knowledge, vSkipper is the first system to realize serving-efficiency gains from per-token interior layer skipping within a modern LLM serving engine. The code is open-sourced as an SGLang fork at https://github.com/AKafakA/sglang-vskipper/tree/vskipper-ref
Comments25 pages (10-pages main body), 6 figures