CommentsWe found a critical flaw in the prompt complexity metric, which affects the 2D curriculum grid construction and leads to potentially invalid comparisons. Since this undermines our main conclusions, we are withdrawing the paper and will revise the methodology before resubmission
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning
RolloutPipe: 分离式在线策略LLM强化学习中的流水线化Rollout与训练重叠
Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Southern University of Science and Technology(南方科技大学)
专题命中
后训练与偏好优化
:LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)
Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning
以合适的速度学习:自适应数据调度改进大语言模型强化学习
Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang, Xiuyi Lou, Hoang Anh Duy Le, Oren Gal, Alexander S. Szalay, Zhaozhuo Xu, Guanchu Wang, Vladimir Braverman
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
Rice University(莱斯大学)
;
University of Haifa(海法大学)
;
Workato
;
University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
专题命中
后训练与偏好优化
:LLM(title,summary_cn);large language model(abstract);language model(abstract);post-training(abstract)
BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM
BayLing-Duplex: 单一自回归LLM的原生全双工语音对话
Qingkai Fang, Shoutao Guo, Yang Feng
机构
*
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院计算技术研究所智能信息处理重点实验室)
;
Key Laboratory of AI Safety, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)