RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
RC-GRPO:基于奖励的组相对策略优化用于多轮工具调用智能体
Haitian Zhong, Jixiu Zhai, Lei Song, Jiang Bian, Qiang Liu, Tieniu Tan
机构
*
New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(模式识别新实验室、多模态人工智能系统国家重点实验室、自动化研究所,中国科学院)
;
School of Mathematics and Statistics(数学与统计学学院)
;
Statistics, Lanzhou University(统计学,兰州大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Microsoft Research(微软研究院)
;
Nanjing University(南京大学)
;
Zhongguancun Academy(中关村学院)
专题命中
后训练与偏好优化
:large language model(abstract);language model(abstract);SFT(abstract);分类 cs.CL、cs.AI
CommentsAccepted for publication in Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT '25), to appear June 2025
机构
*
Shanghai University(上海大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Artificial Intelligence Incubation and Innovation Institute, Fudan University(复旦大学人工智能孵化与创新研究院)
;
Shanghai Academy of AI for Science(上海人工智能科学研究院)
专题命中
后训练与偏好优化
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG