Improving LLM Safety Alignment with Dual-Objective Optimization
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(abstract);jailbreak(abstract)
Comments ICML 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(abstract);jailbreak(abstract)
Comments ICML 2025
机构 * BAISH | UBA | Apart Research(BAISH | UBA | Apart研究) ; University of São Paulo(圣保罗大学) ; Apart Research(Apart研究) ; Dovetail Research | Apart Research(Dovetail研究 | Apart研究)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Stanford University(斯坦福大学) ; University of Toronto(多伦多大学) ; University of Pennsylvania(宾夕法尼亚大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing(中国人民大学北京校区人工智能学院) ; Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索与推荐工程研究中心) ; Beijing Key Laboratory of Research on Large Models(北京大型模型研究重点实验室)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI
Comments Accepted By NeurIPS 2025