RSPO: Regularized Self-Play Alignment of Large Language Models
机构 * University College London(伦敦大学学院) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments Preprint
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University College London(伦敦大学学院) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments Preprint
机构 * The Hong Kong Polytechnic University(香港理工大学)
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
机构 * Zhejiang University(浙江大学) ; Angelalign Technology Inc.(Angelalign技术公司) ; Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江省医学影像人工智能重点实验室)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments ACL Finding
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science(多媒体信息处理国家重点实验室、计算机科学学院)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted by ACL 2025 Findings
机构 * UC Davis(加州大学戴维斯分校) ; USC(南加州大学) ; UW-Madison(威斯康星大学麦迪逊分校)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments ACL 2025
机构 * Soochow University(苏州大学) ; Zhipu AI(智谱AI) ; Tsinghua University(清华大学) ; Key Laboratory of Data Intelligence and Advanced Computing, Soochow University(苏州大学数据智能与先进计算重点实验室)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to the Findings of ACL 2025
机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) ; Nanyang Technological University(南洋理工大学) ; Peng Cheng Laboratory(鹏城实验室) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; Tencent(腾讯)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.AI、cs.LG
Comments 18 pages, to appear in ACL'25
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments Accepted by ICML 2025
机构 * School of Computer Science, Peking University(北京大学计算机学院)
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments CVPR 2025, project page at https://github.com/google-deepmind/video_comp
专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG
Comments ICLR'25
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments Published as a conference paper at AISTATS 2025
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments Published as a conference paper at ICLR 2025
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI
Comments 14 pages, 3 figures
专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments ICLR 2025
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.CY
Comments ACM CHI 2025
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments 35 pages
Journal ref Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025)
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to ICLR 2025
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG
Comments ICLR 2025
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments ICLR 2025
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG
Comments Original work accepted at NeurIPS Efficient Natural Language and Speech Processing Workshop. PMLR, 2024. Experiments with a larger model from a different family, Llama-30B have been added to the appendix for generalizability
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Coling 2025