Rollplex:面向视觉语言模型后训练的跨阶段GPU空间共享
Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
- HKUST(香港科技大学)
- Alibaba Inc(阿里巴巴公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Rollplex是一种跨阶段GPU空间共享运行时,通过解耦RL后训练的参考与训练阶段、优化内存与并行度,提升VLMs后训练的GPU利用率与训练速度。
AI中文摘要:
视觉语言模型(Vision-language models, VLMs)使具身智能体能够基于视觉观测和语言指令进行推理与行动。强化学习(Reinforcement learning, RL)后训练通过任务反馈增强这些能力,但当前的在线策略RL运行时会以严格的串行阶段执行rollout(试 rollout)、参考评分和actor(策略网络)训练。虽然这种方式对纯文本RL有效,但对VLMs而言却存在资源浪费——处理密集视频输入和提示前缀会占用每个阶段的大量时间。由于前缀处理与生成的响应无关,它可以与rollout解码并行运行,且不会破坏同步在线策略语义,同时还能避免GPU算力的闲置。本文提出Rollplex,一种将参考阶段和训练阶段解耦、并将前缀计算移至rollout解码窗口的运行时。实现该调度不仅需要并发内核启动:直接部署Qwen2.5-VL-32B模型时每个GPU需要约165GiB内存,而rollout和训练偏好不同的张量并行(Tensor-Parallel, TP)度和权重布局。Rollplex通过两种机制解决这些约束:1. 感知阶段的内存管理,根据生产者-消费者生命周期控制高带宽内存(HBM)驻留;2. 感知并行度的权重共享,在不同TP度间对布局兼容的张量使用同一物理存储,仅重构不兼容的张量,避免完整的第二actor副本。在32块H800 GPU上,Rollplex在相同GPU预算下相较于串行部署实现了1.23倍至1.30倍的加速,相较于分散部署实现了1.57倍至2.24倍的加速,同时保留了同步RL更新机制。
英文摘要:
Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While effective for text-only RL, this phase-granular execution is wasteful for VLMs, where processing dense video inputs and prompt prefixes occupies a large fraction of each phase. Because prefix processing is independent of the generated response, it can be run alongside rollout decoding, which leaves GPU compute capacity underutilized, without breaking synchronous on-policy semantics. We present Rollplex, a runtime that decomposes the reference and training phase and moves the prefix computation into the rollout decode window. Realizing this schedule requires more than concurrent kernel launches: naive colocation of Qwen2.5-VL-32\,B requires roughly 165\,GiB per GPU, while rollout and training prefer different tensor-parallel (TP) degrees and weight layouts. Rollplex addresses these constraints with two mechanisms. Phase-aware memory management controls HBM residency according to producer--consumer lifetimes. Parallelism-aware weight sharing uses the same physical storage for layout-compatible tensors across distinct TP degrees and reconstructs only incompatible tensors, avoiding a complete second actor copy. On 32 H800 GPUs, Rollplex achieves $1.23\times$--$1.30\times$ speedup over serial colocation and $1.57\times$--$2.24\times$ over disaggregation under the same GPU budget, while preserving the synchronous RL update.