arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22788cs.AIcs.LG

TailSieve:面向大语言模型(LLM)rollout的部分rollout引导式长尾路由

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Ba… 展开作者

Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, Mingchen Wan

首次发表
浏览论文内容

中文总结 AI 辅助

TailSieve是面向LLM rollout的部分rollout引导式框架,通过联合控制长尾路由与副本分配,仅路由即可实现最高1.67倍加速,结合专用推测解码可达2.59倍加速,同时保留策略内生成并避免长度偏差。

中文摘要 AI 辅助

大规模rollout已成为现代LLM系统的核心组件,涵盖强化学习(RL)后训练、策略蒸馏(OPD)以及采样密集型评估流程。与针对请求级延迟和吞吐量优化的在线服务不同,少数长尾生成内容会主导整个rollout步骤的端到端总时长。实际中,rollout请求常被均匀路由至各副本,这可能导致极长的生成长度被置于高并发解码批次中。为解决该问题,本文提出TailSieve,这是一种部分rollout引导的框架,可联合控制LLM rollout的长尾路由与副本分配。在完成长度已知的理想化场景中,研究表明长尾场景下总时长最优的路由需结合长尾隔离与负载均衡,且简单的top-k策略能近似该离线最优值。利用长尾提示在策略更新过程中往往仍保持长尾分布的观察结果,TailSieve使用部分rollout作为无训练信号来识别候选长尾组。随后,分层控制器结合收集的响应-工作历史与测得的并发-吞吐量模型,联合调整隔离组的数量以及长尾池与批量池之间的副本分配。TailSieve仅通过路由即可实现比均匀组路由最高1.67倍的加速;由此产生的低并发长尾池还可支持使用MTP或DFlash的路由专用推测解码,实现比均匀路由最高2.59倍的加速。所选提示会在当前策略下重新生成,保留策略内生成,且稳态下可避免额外路由引发的长度偏差。

英文摘要

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.

发表机构

  • Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Zhejiang University(浙江大学)
  • University of Science and Technology of China(中国科学技术大学)
  • National University of Singapore(新加坡国立大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑