Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
Sheng Ouyang, Yulan Hu, Ge Chen, Qingyang Li, Fuzheng Zhang, Yong Liu
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
;
Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室)
;
Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部)
;
Kuaishou Technology(快手科技)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
后训练与偏好优化
:RLHF(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
Sureyya Akin, Shruti T. Tiwari, Ram Bhattacharya, Sagar A. Raman, Kiran Mohanty, Sita Krishnan
专题命中
长上下文与记忆
:large language model(title,abstract);language model(title,abstract)
CommentsWe have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript
From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
Siyu Xia, Zekun Xu, Jiajun Chai, Wentian Fan, Yan Song, Xiaohan Wang, Guojun Yin, Wei Lin, Haifeng Zhang, Jun Wang
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Meituan(美团)
;
Nanjing University of Posts and Telecommunications(南京邮电大学)
;
AI Centre, Department of Computer Science, University College London(伦敦大学学院计算机科学系人工智能中心)
专题命中
长上下文与记忆
:LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)
Continuous Subspace Optimization for Continual Learning
Quan Cheng, Yuanyu Wan, Lingyu Wu, Chenping Hou, Lijun Zhang
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学)
;
School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)
;
School of Software Technology, Zhejiang University(软件学院,浙江大学)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou, China(杭州高新技术区(滨江)区块链与数据安全研究院,杭州,中国)
;
College of Science, National University of Defense Technology(科学学院,国防科技大学)
机构
*
Singapore Management University(新加坡管理大学)
;
University of Rochester(罗切斯特大学)
;
University College London(伦敦大学学院)
;
National University of Singapore(新加坡国立大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Stanford University(斯坦福大学)
LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
Xiaohan Zhang, Yan Ding, Yohei Hayamizu, Zainab Altaweel, Yifeng Zhu, Yuke Zhu, Peter Stone, Chris Paxton, Shiqi Zhang
机构
*
The State University of New York at Binghamton(纽约州立大学布林顿分校)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Sony AI(索尼人工智能)
;
Hello Robot(Hello Robot公司)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
OneStar Robotics(OneStar机器人)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);LLM(title)
Journal refThe International Journal of Robotics Research, 2025, Vol. 0(0), pp. 1-19
机构
*
University of Milan-Bicocca(米兰-比科卡大学)
;
University of Naples Federico II(那不勒斯费德里科二世大学)
;
Oversonic Robotics(Oversonic机器人公司)
;
University of Essex(埃塞克斯大学)
;
TUM School of Social Sciences and Technology(慕尼黑技术大学社会科学与技术学院)
专题命中
推理与问题求解
:LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)