超越前缀局部性的混合强化学习回滚调度
Scheduling Mixed RL Rollouts Beyond Prefix Locality
浏览论文内容
中文总结 AI 辅助
针对混合RL回滚调度的异构性问题,提出MISA-T策略,在多基准实验中显著提升回滚吞吐量并降低平均轮次时间,同时维持缓存命中率与工作负载混合比例。
中文摘要 AI 辅助
大型语言模型(LLM)的现代强化学习(RL)后训练流程日益将多个领域及反馈范式的回滚工作负载相结合。感知前缀的路由通过缓存复用和负载均衡提升推理效率,但无法控制异构回滚会话对KV缓存容量的竞争。当带可验证奖励的强化学习(RLVR)、人类反馈强化学习(RLHF)及智能体回滚共享异步推理服务时,它们独特的序列结构、交互模式和KV驻留时间会产生截然不同的服务需求。回滚调度必须在不破坏训练器指定的工作负载混合的前提下考虑这种异构性。我们提出MISA-T,一种用于混合回滚服务的路由层准入策略,它结合了自适应会话准入、感知工作负载的KV容量分配以及感知驻留时间的KV核算。在仅回滚的 ablation 实验中,针对Step3.7和Qwen3.6-35B-A3B,MISA-T相比经 sweep 调优的感知缓存vLLM Router分别提升了53.3%和43.6%的回滚吞吐量,同时保持了较高的前缀缓存命中率。在匹配的50轮Step3.7实验中,它将回滚吞吐量提升了35.6%,将平均轮次时间降低了22.8%,同时保持消耗的工作负载混合接近训练器目标,并取得了可比的任务分数。
英文摘要
Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifiable rewards (RLVR), reinforcement learning from human feedback (RLHF), and agentic rollouts share an asynchronous inference service, their distinct sequence structures, interaction patterns, and KV-residency times create substantially different serving demands. Rollout scheduling must account for this heterogeneity without distorting the workload mixture specified by the trainer. We present MISA-T, a routing-layer admission policy for mixed rollout serving. MISA-T combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting. In rollout-only ablations on Step3.7 and Qwen3.6-35B-A3B, MISA-T improves rollout throughput over a sweep-tuned cache-aware vLLM Router by 53.3% and 43.6%, respectively, while maintaining high prefix-cache hit rates. In a matched 50-iteration Step3.7 experiment, it increases rollout throughput by 35.6% and reduces mean iteration time by 22.8%, while keeping the consumed workload mixture close to the trainer target and achieving comparable task scores.
发表机构
- State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
- StepFun(阶跃星辰)
机构由 AI 辅助整理,请以论文原文为准。