机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Baidu Inc.(百度公司)
专题命中
后训练与偏好优化
:large language model(title,abstract);language model(title,abstract);preference optimization(title,abstract);SFT(abstract)
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
;
Joint corresponding authors(联合通讯作者)
专题命中
后训练与偏好优化
:large language model(title,abstract);language model(title,abstract);RLHF(title,abstract);LLM(abstract)
CommentsAccepted for publication in the Journal of the American Statistical Association
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado
机构
*
Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国)
;
Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国)
;
Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)
机构
*
School of Software Engineering, Xi’an Jiaotong University(软件工程学院,西安交通大学)
;
School of Computer Science and Technology and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(计算机科学与技术学院和教育部智能网络与网络安全重点实验室,西安交通大学)
;
School of Automation, Xi’an Jiaotong University(自动化学院,西安交通大学)
;
College of Artificial Intelligence, Xi’an Jiaotong University(人工智能学院,西安交通大学)
;
School of Mathematics and Statistics and Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学院和教育部智能网络与网络安全重点实验室,西安交通大学)
机构
*
the Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
the institute of information engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
the College of Computer, National University of Defense Technology(国防科技大学计算机学院)
;
the University of Sydney(悉尼大学)
专题命中
长上下文与记忆
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
CommentsThis paper has been accepted by Neurocomputing
机构
*
School of Information Sciences, University of Illinois Urbana-Champaign(信息科学学院,伊利诺伊大学厄巴纳-香槟分校)
;
Department of Computer Science, University of Oxford(计算机科学系,牛津大学)
专题命中
长上下文与记忆
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks
Ye Ye
机构
*
Ye Ye(独立研究者)
专题命中
长上下文与记忆
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments14 pages, 5 figures. Preprint prepared for future submission. Includes implementation and token-efficiency analysis. Code at https://github.com/biubiutomato/TME-Agent