Katherine M. Collins, Cedegao E. Zhang, Graham Todd, Lance Ying, Mauricio Barba da Costa, Ryan Liu, Prafull Sharma, Adrian Weller, Ionatan Kuperwajs, Lionel Wong, Joshua B. Tenenbaum, Thomas L. Griffiths
机构
*
University of Cambridge(剑桥大学)
;
MIT(麻省理工学院)
;
Princeton University(普林斯顿大学)
;
NYU(纽约大学)
;
Harvard University(哈佛大学)
;
Stanford University(斯坦福大学)
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
Columbia University(哥伦比亚大学)
;
California State University(加州州立大学)
;
University of Montreal(蒙特利尔大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Rensselaer Polytechnic Institute(莱斯利理工学院)
;
The University of Manchester(曼彻斯特大学)
;
Harvard University(哈佛大学)
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学杭州全球科学与技术创新中心)
;
Department of Chemistry, Fudan University(复旦大学化学系)
机构
*
State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所处理器重点实验室)
;
School of Advanced Interdisciplinary Sciences, CAS(中国科学院高等交叉学科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Microelectronics, CAS(中国科学院微电子研究所)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
RubricEM: 通过规则引导的策略分解实现超越可验证奖励的元强化学习
Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Google Cloud AI Research(谷歌云人工智能研究)
机构
*
School of Computer Science, Guangdong University of Technology(广东技术大学计算机科学学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))
;
Peng Cheng Laboratory(鹏城实验室)
;
College of Science, Shantou University(汕头大学理学院)
Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt
小语言模型帮助解决LLM提示的语义歧义
Zhenzhen Huang, Chaoning Zhang, Fachrina Dewi Puspitasari, Jiaquan Zhang, Yitian Zhou, Shuxu Chen, Yang Yang
机构
*
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Department of Electronic Engineering, Kyung Hee University(庆熙大学电子工程系)
CommentsAccepted to the Adaptive and Learning Agents Workshop (ALA 2026) @ AAMAS 2026. Code is available at github.com/vicgalle/experiential-prompt-optimization-safe
Journal refProc. of the Adaptive and Learning Agents Workshop (ALA 2026) @ AAMAS 2026