Cross-Modal Safety Alignment: Is textual unlearning all you need?
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments Accepted by EMNLP 2024 Findings
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments Accepted by EMNLP 2024 Findings
机构 * The Chinese University of Hong Kong(香港中文大学) ; Fudan University(复旦大学)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2025
机构 * Yale University(耶鲁大学) ; Allen Institute for AI(人工智能研究院)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Sun Yat-sen University(中山大学) ; X-Era AI Lab(X-Era人工智能实验室)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * University of Cambridge(剑桥大学) ; Apta
专题命中 偏好对齐 :alignment(abstract);分类 cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL
Comments 5 Pages, 4 Figures, 4 Tables
Journal ref 39th Conference on Neural Information Processing Systems, 2025, Workshop: Reliable ML from Unreliable Data