A Principled Loss Function for Direct Language Model Alignment
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * University of Science and Technology of China(中国科学技术大学) ; Mach Drive(马车驱动)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract)
机构 * Arizona State University(亚利桑那州立大学)
专题命中 偏好对齐 :alignment(title,abstract)
Comments EMNLP 2025 (Main)
机构 * LLM Department, Tencent(腾讯大模型部门) ; HunYuan Infra Team(文心一言基础设施团队) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Work in progress
机构 * Preferred Networks Inc(Preferred Networks公司)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Peking University(北京大学) ; Tsinghua University(清华大学) ; Mila - Québec AI Institute(魁北克人工智能研究所)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
Comments NeurIPS 2025. Code: https://github.com/Gen-Verse/HermesFlow
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI
Comments 31 pages, 16 figures, 12 tables
机构 * Indian Institute of Technology Patna(印度帕纳杰大学) ; CRISIL LTD(CRISIL公司)
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI