Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts
机构 * CyberAgent / Tokyo, Japan(CyberAgent)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP Findings, 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * CyberAgent / Tokyo, Japan(CyberAgent)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP Findings, 2025
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) ; Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG
Comments 19pages, 13 figures, 11 tables
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
Comments 38 pages, 31 figures
机构 * Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系) ; School of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学学院) ; Department of Psychology, University of Warwick, UK(沃里克大学心理学系)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL
Comments EMNLP 2025 Findings camera-ready, 9+7 pages