Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
全栈对齐:通过厚价值模型对齐人工智能与机构
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe, Oliver Klingefjord, Vincent Wang-Mascianica, Matija Franklin, Ryan Othniel Kearns, Ellie Hain, Atrisha Sarkar, Michiel Bakker, Fazl Barez, David Duvenaud, Jakob Foerster, Iason Gabriel, Joseph Gubbels, Bryce Goodman, Andreas Haupt, Jobst Heitzig, Julian Jara-Ettinger, Atoosa Kasirzadeh, James Ravi Kirkpatrick, Andrew Koh, W. Bradley Knox, Philipp Koralus, Joel Lehman, Sydney Levine, Samuele Marro, Manon Revel, Toby Shorin, Morgan Sutherland, Michael Henry Tessler, Ivan Vendrov, James Wilken-Smith
机构
*
Meaning Alignment Institute(意义对齐研究所)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University College London(伦敦大学学院)
;
University of Oxford(牛津大学)
;
Western University(西方大学)
;
University of Toronto(多伦多大学)
;
McGill University(麦吉尔大学)
;
Stanford University(斯坦福大学)
;
Potsdam Institute for Climate Impact Research(波茨坦气候影响研究所)
;
Yale University(耶鲁大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
UT Austin(德克萨斯大学奥斯汀分校)
;
New York University(纽约大学)
;
Harvard University(哈佛大学)
;
Midjourney Core contributor(Midjourney核心贡献者)
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
大语言模型奖励学习中异质人类反馈的不确定性量化
Pangpang Liu, Junwei Lu, Will Wei Sun
机构
*
Department of Biostatistics, Yale University(耶鲁大学生物统计学系)
;
Department of Biostatistics, Harvard University(哈佛大学生物统计学系)
;
Department of Quantitative Methods, Purdue University(普渡大学定量方法系)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
轨迹平衡与异步性:解耦探索与学习以实现快速、可扩展的LLM后训练
Brian Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan Obando-Ceron, Yoshua Bengio, Bhavya Kailkhura
机构
*
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
;
Mila – Quebec AI Institute(魁北克AI研究院)
;
Université de Montréal(蒙特利尔大学)
;
KAIST(韩国科学技术院)
;
CIFAR Fellow