What are Key Factors for Updates in RL for LLM Reasoning?
RL提升LLM推理能力的关键更新因素是什么?
Peidong Wang, Demi Wang, Xufang Luo, Jiahang Xu, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Microsoft Research(微软研究院)
;
Carnegie Mellon University(卡内基梅隆大学)
机构
*
Indian Institute of Technology Bombay(印度理工学院班加罗尔)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
Centre for Machine Intelligence and Data Science(机器智能与数据科学中心)
;
Microsoft Research India(微软印度研究院)
;
Microsoft India(微软印度)
Commentsan extended version of the ICLR 2026 paper (added a paragraph in Sec 3.2 about short-sided KL-Shampoo as scaled Muon when momentum is disabled)
CommentsProceedings of the 39th Annual Conference on Neural Information Processing Systems, ARLET Workshop (Aligning Reinforcement Learning Experimentalists and Theorists)
Journal refTransactions on Machine Learning Research, Vol. 2026, June 2026
Measuring Human Contribution in AI-Assisted Content Generation
衡量AI辅助内容生成中的人类贡献
Yueqi Xie, Tao Qi, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang, Qian Ding, Yu Xie, Xing Xie, Fangzhao Wu
机构
*
Princeton University(普林斯顿大学)
;
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
The University of Hong Kong(香港大学)
;
Microsoft Research Asia(微软亚洲研究院)
;
Peking University(北京大学)