HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
HISR: 基于 hindsight 信息的分段过程奖励用于多轮代理强化学习
Zhicong Lu, Zichuan Lin, Wei Jia, Changyuan Tian, Deheng Ye, Peiguang Li, Li Jin, Nayu Liu, Guangluan Xu, Wei Feng
机构
*
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tencent Hunyuan(腾讯文脉)
;
School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
School of ECE, Peking University(北京大学电子工程学院)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Tencent Youtu Lab(腾讯优图实验室)
Enhancing Lexicon-Based Text Embeddings with Large Language Models
通过大语言模型增强基于词典的文本嵌入
Yibin Lei, Tao Shen, Yu Cao, Andrew Yates
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
University of Technology Sydney(技术悉尼大学)
;
Tencent IEG(腾讯IEG)
;
Johns Hopkins University, HLTCOE(约翰霍普金斯大学,HLTCOE)
机构
*
School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)
;
ARC Lab, Tencent PCG(腾讯PCG ARC实验室)
;
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology(广东省超高清沉浸媒体技术重点实验室)
;
GVC Lab, Great Bay University(Great Bay大学GVC实验室)
;
The Chinese University of Hong Kong(香港中文大学)