机构
*
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
DP Technology(DP科技)
;
Meituan(美团)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
强化学习中大语言模型的进化系统提示学习
Lunjun Zhang, Ryan Chen, Bradly C. Stadie
机构
*
Department of Computer Science, University of Toronto(多伦多大学计算机科学系)
;
Department of Statistics(统计学系)
;
Data Science, Northwestern University(数据科学,西北大学)
;
Bridgewater AIA Labs(布里奇沃特AIA实验室)
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Tsinghua University(清华大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Science and Technology of China(中国科学技术大学)
Simi Job, Xiaohui Tao, Taotao Cai, Haoran Xie, Jianming Yong, Xin Wang
机构
*
School of Mathematics, Physics, and Computing, University of Southern Queensland(数学、物理与计算学院,南方昆士兰大学)
;
School of Data Science, Lingnan University(数据科学学院,岭南大学)
;
School of Business, University of Southern Queensland(商学院,南方昆士兰大学)
;
Schulich School of Engineering, University of Calgary(Schulich工程学院,卡尔加里大学)
机构
*
School of Physics, Peking University(物理系,北京大学)
;
School of Electronics Engineering and Computer Science, Peking University(电子工程与计算机科学系,北京大学)
;
Center for High Energy Physics, Peking University(高能物理中心,北京大学)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
TAROT: 为基于大语言模型的代码生成设计的测试驱动和能力适应课程强化微调
Chansung Park, Juyong Jiang, Fan Wang, Sayak Paul, Jiasi Shen, Jing Tang, Jianguo Li
机构
*
Electronics and Telecommunications Research Institute(电信研究所)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Hugging Face
;
Ant Group(蚂蚁集团)
Large Language Model (LLM)-enabled Reinforcement Learning for Wireless Network Optimization
基于大语言模型的强化学习用于无线网络优化
Jie Zheng, Ruichen Zhang, Dusit Niyato, Haijun Zhang, Jiacheng Wang, Hongyang Du, Jiawen Kang, Zehui Xiong
机构
*
State-Province Joint Engineering and Research Center of Advanced Networking and Intelligent Information Services, College of Computer Science, Northwest University(高级网络与智能信息服务省-市联合工程与研究中心,计算机科学学院,西北大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
Institute of Artificial Intelligence, University of Science and Technology Beijing(人工智能研究院,北京科技大学)
;
Department of Electrical and Electronic Engineering, the University of Hong Kong(电子与电气工程系,香港大学)
;
Automation of School, Guangdong University of Technology(自动化学院,广东工业大学)
;
Queen’s University Belfast(贝尔法斯特女王大学)
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))
;
Southeast University(东南大学)
;
University of Tsukuba(茨口大学)
Learning to Self-Verify Makes Language Models Better Reasoners
学习自我验证使语言模型更擅长推理
Yuxin Chen, Yu Wang, Yi Zhang, Ziang Ye, Zhengzhou Cai, Yaorui Shi, Qi Gu, Hui Su, Xunliang Cai, Xiang Wang, An Zhang, Tat-Seng Chua
机构
*
National University of Singapore(新加坡国立大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
GraphAgents: Knowledge Graph-Guided Agentic AI for Cross-Domain Materials Design
GraphAgents: 基于知识图谱的多智能体AI用于跨领域材料设计
Isabella A. Stewart, Tarjei Paule Hage, Yu-Chuan Hsu, Markus J. Buehler
机构
*
Department of Civil and Environmental Engineering Massachusetts Institute of Technology(土木与环境工程系 马萨诸塞理工学院)
;
Department of Mechanical Engineering Massachusetts Institute of Technology(机械工程系 马萨诸塞理工学院)
;
Department of Civil and Environmental Engineering Department of Mechanical Engineering Schwarzman College of Computing Massachusetts Institute of Technology(土木与环境工程系 机械工程系 斯沃茨曼计算学院 马萨诸塞理工学院)
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
并非所有层都需要调节:选择性层恢复恢复多样性
Bowen Zhang, Meiyi Wang, Harold Soh
机构
*
Department of Computer Science, National University of Singapore, Singapore(新加坡国立大学计算机科学系)
;
Smart Systems Institute, National University of Singapore, Singapore, Singapore(新加坡国立大学智能系统研究所)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
CompactRAG: 减少多跳问答中的LLM调用和令牌开销
Hao Yang, Zhiyu Yang, Xupeng Zhang, Wei Wei, Yunjie Zhang, Lin Yang
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
;
Erik Jonsson School of Engineering and Computer Science, University of Texas at Dallas(埃里克·乔纳森工程与计算机科学学院,德克萨斯大学达拉斯分校)
;
Isoftstone Information Technology (Group) Co.,Ltd.(伊软石信息技术(集团)有限公司)
;
College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学)
;
School of Electronic Information, Central South University(电子信息学院,中南大学)
机构
*
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)
;
Tsinghua University(清华大学)
;
AutoGame Research(AutoGame研究)
;
Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳人工智能与机器人研究院)