机构
*
Meta Superintelligence Labs(Meta超智能实验室)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Arizona State University(亚利桑那州立大学)
;
University of Southern California(南加州大学)
How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
采样如何塑造大语言模型对齐:从单次最优到迭代动态
Yurong Chen, Yu He, Michael I. Jordan, Fan Yao
机构
*
Inria(法国国家信息与自动化研究所)
;
École Normale Supérieure(巴黎高等师范学校)
;
PSL Research University(巴黎综合理工研究所)
;
Northwestern University(西北大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Stackelberg 自注释:一种鲁棒的数据高效 LLM 对齐方法
Xu Chu, Zhixin Zhang, Tianyu Jia, Yujie Jin
机构
*
Key Laboratory of High Confidence Software Technologies, Ministry of Education(高可信软件技术重点实验室,教育部)
;
Center on Frontiers of Computing Studies, Peking University(计算前沿研究中心,北京大学)
;
School of Computer Science, Peking University(计算机学院,北京大学)
DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
DA-DPO:面向减少多模态大语言模型幻觉的高效难度感知偏好优化
Longtian Qiu, Shan Ning, Chuyu Zhang, Jiaxuan Sun, Xuming He
机构
*
ShanghaiTech University(上海科技大学)
;
Lingang Laboratory(灵冈实验室)
;
Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)
Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
全栈对齐:通过厚价值模型对齐人工智能与机构
Joe Edelman, Tan Zhi-Xuan, Ryan Lowe, Oliver Klingefjord, Vincent Wang-Mascianica, Matija Franklin, Ryan Othniel Kearns, Ellie Hain, Atrisha Sarkar, Michiel Bakker, Fazl Barez, David Duvenaud, Jakob Foerster, Iason Gabriel, Joseph Gubbels, Bryce Goodman, Andreas Haupt, Jobst Heitzig, Julian Jara-Ettinger, Atoosa Kasirzadeh, James Ravi Kirkpatrick, Andrew Koh, W. Bradley Knox, Philipp Koralus, Joel Lehman, Sydney Levine, Samuele Marro, Manon Revel, Toby Shorin, Morgan Sutherland, Michael Henry Tessler, Ivan Vendrov, James Wilken-Smith
机构
*
Meaning Alignment Institute(意义对齐研究所)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University College London(伦敦大学学院)
;
University of Oxford(牛津大学)
;
Western University(西方大学)
;
University of Toronto(多伦多大学)
;
McGill University(麦吉尔大学)
;
Stanford University(斯坦福大学)
;
Potsdam Institute for Climate Impact Research(波茨坦气候影响研究所)
;
Yale University(耶鲁大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
UT Austin(德克萨斯大学奥斯汀分校)
;
New York University(纽约大学)
;
Harvard University(哈佛大学)
;
Midjourney Core contributor(Midjourney核心贡献者)
机构
*
Department of Informatics, University of Sussex(信息学院,苏塞克斯大学)
;
Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学)
;
Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学)
;
Department of Computer Science, Purdue University(计算机科学系,普渡大学)
;
Department of Computer Science, Emory University(计算机科学系,埃默里大学)
;
AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团)
;
Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)
Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Fei Long, Rundi Zhai, Dan Li, Yanyu Ren, Tianfeng Liu, Hongtao Xie, Ce Yang, Xuhui Cai
机构
*
Zhongguancun Laboratory(中关村实验室)
;
Tsinghua University(清华大学)
;
China Mobile Communications Group Co., Ltd.(中国移动通信集团有限公司)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术实验室,南京大学)
;
School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)
专题命中
偏好对齐
:RLHF(title,abstract);分类 cs.LG
CommentsNeurIPS 2025; The first two authors contributed equally
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
Huahui Yi, Kun Wang, Qiankun Li, Miao Yu, Liang Lin, Gongli Xi, Hao Wu, Xuming Hu, Kang Li, Yang Liu
机构
*
West China Biomedical Big Data Center, West China Hospital, SCU(西昌生物医学大数据中心、西昌医院、SCU)
;
NTU
;
USTC
;
TeleAI, China Telecom(TeleAI、中国电信)
;
BUPT
;
Tsinghua University(清华大学)
;
HKUST(Guangzhou)(HKUST(广州))
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
Shivanshu Shekhar, Shreyas Singh, Tong Zhang
机构
*
Siebel School of Computing and Data Science(计算与数据科学学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Fractal AI Research(Fractal AI研究院)