机构
*
Google DeepMind(谷歌DeepMind)
;
Indian Institute of Technology Bombay(印度理工学院孟买分校)
;
International Centre for Theoretical Sciences, Tata Institute of Fundamental Research(理论科学国际中心, Tata 基础研究机构)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
REAL: 面向LLM评判的回归感知强化学习
Yasi Zhang, Tianyu Chen, Mingyuan Zhou, Oscar Leong, Ying Nian Wu, Michal Lukasik
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
The University of Texas at Austin(得克萨斯大学奥斯汀分校)
;
Google Research Now at Google DeepMind(谷歌研究 现在在谷歌深Mind)
Discovering Differences in Strategic Behavior Between Humans and LLMs
发现人类与LLM在战略行为上的差异
Caroline Wang, Daniel Kasenberg, Kim Stachenfeld, Pablo Samuel Castro
机构
*
Department of Computer Science, University of Texas at Austin. Work performed as a student researcher at Google DeepMind(德克萨斯大学计算机科学系。作为谷歌DeepMind的学生研究员进行的工作)
;
Google DeepMind(谷歌DeepMind)