What are Key Factors for Updates in RL for LLM Reasoning?
RL提升LLM推理能力的关键更新因素是什么?
Peidong Wang, Demi Wang, Xufang Luo, Jiahang Xu, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Microsoft Research(微软研究院)
;
Carnegie Mellon University(卡内基梅隆大学)
Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation
先看后思:解耦感知与推理以实现抗捷径的多模态在策略自蒸馏
Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han
机构
*
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning
ARMOR-MAD:大语言模型推理中异构多智能体辩论的自适应路由
Fuqiang Niu, Bowen Zhang
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络空间安全学院)
;
School of Artificial Intelligence, Shenzhen Technology University(深圳技术大学人工智能学院)
机构
*
University of Georgia(佐治亚大学)
;
Tencent AI Lab(腾讯AI实验室)
;
The Education University of Hong Kong(香港教育大学)
;
The Hong Kong Polytechnic University(香港理工大学)
PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR
PAEC:面向RLVR中LLM推理的位置感知熵校准
Shumeng Yang, Yisu Liu, Jiayi Zheng, Zhaohui Yang, Linjing Li
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning
好的推理产生好的示范:通过上下文强化学习进行隐式推理质量监督
Tiehua Mei, Minxuan Lv, Leiyu Pan, Zhenpeng Su, Hongru Hou, Hengrui Chen, Ao Xu, Deqing Yang
机构
*
School of Data Science, Fudan University(复旦大学数据科学学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院)
Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization
协同进化智能体架构与可解释推理用于自动化优化
Jiahao Huang, Peilan Xu, Xiaoya Nan, Wenjian Luo
机构
*
School of Artificial Intelligence, Nanjing University of Information Science and Technology(南京信息工程大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies, Institute of Cyberspace Security, School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
机构
*
Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China(香港城市大学计算机科学系)
;
Harbin Institute of Technology, Harbin, China(哈尔滨工业大学)
CommentsarXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission. Author list and submitter name redacted due to disputed authorship