并行与分布式推理语言模型的性能基础
Performance Foundations of Parallel & Distributed Reasoning Language Models
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对RL-for-LLM范式,分析主流后训练算法框架,构建RLM并行策略分类体系,提炼实用指南并概述开放方向,以推动开发高性能可扩展的高性价比RLM。
AI中文摘要:
带可验证奖励的强化学习(RLVR)及其他类RL的后训练范式已被用于将大语言模型(LLM)与推理标准对齐。由此产生的近期推理语言模型(RLM),如DeepSeek-R1、o3和Kimi k1.5,表明这类类RL的后训练(“RL-for-LLMs”)可大幅提升思维链推理、长程规划和自我修正能力。然而,这些系统的计算开销巨大:最先进的RLM训练需要数百万GPU小时,且采用紧密耦合的多模型流水线,对现代硬件的压力远超传统有监督LLM训练,这使得RLM训练既是算法问题,也是并行与分布式系统问题。本研究为推动开发高性能、可扩展且高性价比的RLM,首先系统化梳理RL-for-LLM范式,并对主流后训练算法框架——近端策略优化(PPO)、组相对策略优化(GRPO)及其变体——开展以计算为中心的分析;其次,构建RL-for-LLM的模型内与模型间并行策略分类体系,涵盖传统技术(数据并行、张量并行、流水线并行、序列并行、上下文并行及专家并行),以及面向多模型RLM训练的新型并行形式与优化技术,例如拆分放置、阶段融合、混合并行与异步执行;本研究利用并行计算的工作深度模型,使该分类体系及其见解兼具严谨性与可移植性;最后,分析现有RLM框架,提炼实用指南,并概述构建可扩展、快速且高性价比RLM的开放研究方向。
英文摘要:
Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5 show that such RL-style post-training ("RL-for-LLMs") can substantially improve chain-of-thought reasoning, long-horizon planning, and self-correction. However, the computational footprint of these systems is massive: state-of-the-art RLM training requires millions of GPU-hours and tightly coupled multi-model pipelines that stress modern hardware far beyond classical supervised LLM training. This makes RLM training as much a parallel and distributed systems problem as an algorithmic one. In this work, to facilitate developing RLMs that are simultaneously high-performance, scalable, and cost-effective, we first systematize the RL-for-LLM paradigm and provide a compute-centric analysis of prominent post-training algorithmic frameworks: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), as well as their variants. Second, we develop a taxonomy of intra- and inter-model parallelism strategies for RL-for-LLMs, covering both traditional techniques (data, tensor, pipeline, sequence, context, and expert parallelism) as well as novel forms of parallelism and optimization techniques for multi-model RLM training, for example disaggregated placement, stage fusion, hybrid parallelism, and asynchronous execution. We harness the work-depth model of parallel computing to make our taxonomy and its insights rigorous and portable. Finally, we analyze existing RLM frameworks and we distill practical guidelines and outline open research directions for building scalable, fast, and cost-effective RLMs.