The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
机构 * MIT(麻省理工学院)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * MIT(麻省理工学院)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
机构 * Department of Statistics(统计系) ; LSE London, UK(伦敦大学学院) ; Department of Mathematics(数学系) ; Tsinghua University(清华大学) ; School of Design LCC, UAL London, UK(伦敦艺术大学设计学院) ; Department of Engineering Science(工程科学系) ; University of Oxford(牛津大学)
专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG
Comments Accepted to NeurIPS 2025
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL
Comments 18 pages, 6 figures, accepted to Findings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) ; Independent Researcher(独立研究者)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Tsinghua University(清华大学) ; Shanghai Qi Zhi Institute(上海启智研究院) ; Harbin Institute of Technology(哈尔滨工业大学) ; Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团) ; Peng Cheng Laboratory(鹏城实验室) ; National University of Singapore(新加坡国立大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL;RLHF(comments)
Comments Project Website: https://github.com/RLHF-V/RLAIF-V