InfAlign: Inference-aware language model alignment
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract)
机构 * Peking University(北京大学) ; Tsinghua University(清华大学)
专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG
Comments 10 pages, 9 figures, Under review as a full paper at AAAI 2026. A preliminary version is under review at the NeurIPS 2025 Workshop on Reliable ML from Unreliable Data
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI
机构 * University of Ljubljana, Faculty of Computer and Information Science(卢布尔雅那大学计算机与信息科学系)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL
Comments Paper with individual presentation at LUHME workshop at ECAI 2025
专题命中 偏好对齐 :alignment(abstract)
Comments Accepted by The 38th Annual ACM Symposium on User Interface Software and Technology (UIST Adjunct '25), September 28-October 1, 2025, Busan, Republic of Korea