arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向推理大语言模型的强化学习中的Rollout效率:分类与未来方向

Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

Niloofar Gholipour, Marcos Assuncao, Gursimran Singh, Timothy Yu, Rajkumar Buyya, Julien Gascon-Samson, Zhenan Fan, Yong Zhang, Xiaojie Xu, Yaqiang Yao, Xiaolong Bai

arXiv 2609.25463首次发表:更新:

发表机构

École de technologie supérieure, Univ. of Québec; Huawei Technologies; The Univ. of Melbourne(魁北克大学高等技术学院; 华为技术有限公司; 墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述系统分类了面向推理大语言模型强化学习中的rollout效率研究,从机制和瓶颈角度分析技术,并指出评估空白与未来方向。

AI 中文摘要

面向推理的强化学习使大语言模型能够解决数学、编程及其他多步任务,但将训练成本的很大一部分转移到了rollout(轨迹生成)上,即生成用于策略更新的轨迹。因此,高效的rollout机制对于降低这一成本,同时保持训练数据的新鲜度、一致性和统计有效性至关重要。本综述对面向推理的强化学习中rollout效率的近期研究进行了系统分类,从机制和瓶颈两个角度对现有方法进行分类。基于该分类,我们分析了不同技术族如何解决rollout效率低下的不同根源,考察了组合这些技术的机遇与潜在冲突,指出了效率提升评估与报告中的空白,并讨论了开放挑战和未来研究方向。

英文摘要

Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, where trajectories are generated for policy updates. Efficient rollout mechanisms are therefore essential to reduce this cost while maintaining the freshness, consistency, and statistical validity of training data. This survey provides a systematic taxonomy of recent research on rollout efficiency for reasoning-oriented reinforcement learning, classifying existing approaches from both mechanism and bottleneck perspectives. Based on this taxonomy, we analyze how different technique families address distinct sources of rollout inefficiency, examine opportunities and potential conflicts for combining them, identify gaps in the evaluation and reporting of efficiency gains, and discuss open challenges and future research directions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑