TailSieve:面向大语言模型(LLM)rollout的部分rollout引导式长尾路由
TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
浏览论文内容
中文总结 AI 辅助
TailSieve是面向LLM rollout的部分rollout引导式框架,通过联合控制长尾路由与副本分配,仅路由即可实现最高1.67倍加速,结合专用推测解码可达2.59倍加速,同时保留策略内生成并避免长度偏差。
中文摘要 AI 辅助
大规模rollout已成为现代LLM系统的核心组件,涵盖强化学习(RL)后训练、策略蒸馏(OPD)以及采样密集型评估流程。与针对请求级延迟和吞吐量优化的在线服务不同,少数长尾生成内容会主导整个rollout步骤的端到端总时长。实际中,rollout请求常被均匀路由至各副本,这可能导致极长的生成长度被置于高并发解码批次中。为解决该问题,本文提出TailSieve,这是一种部分rollout引导的框架,可联合控制LLM rollout的长尾路由与副本分配。在完成长度已知的理想化场景中,研究表明长尾场景下总时长最优的路由需结合长尾隔离与负载均衡,且简单的top-k策略能近似该离线最优值。利用长尾提示在策略更新过程中往往仍保持长尾分布的观察结果,TailSieve使用部分rollout作为无训练信号来识别候选长尾组。随后,分层控制器结合收集的响应-工作历史与测得的并发-吞吐量模型,联合调整隔离组的数量以及长尾池与批量池之间的副本分配。TailSieve仅通过路由即可实现比均匀组路由最高1.67倍的加速;由此产生的低并发长尾池还可支持使用MTP或DFlash的路由专用推测解码,实现比均匀路由最高2.59倍的加速。所选提示会在当前策略下重新生成,保留策略内生成,且稳态下可避免额外路由引发的长度偏差。
英文摘要
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.
发表机构
- Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
- Carnegie Mellon University(卡内基梅隆大学)
- Zhejiang University(浙江大学)
- University of Science and Technology of China(中国科学技术大学)
- National University of Singapore(新加坡国立大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。