Cascade Reward Sampling for Efficient Decoding-Time Alignment
机构 * Department of Computer Science(计算机科学系)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Department of Computer Science(计算机科学系)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * Florida State University(佛罗里达州立大学)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted for publication in the Proceedings of the 5th Workshop on Bias and Fairness in AI (BIAS 2025) at ECML PKDD
机构 * School of Computer Engineering, Jimei University, Xiamen, 361021, China(厦门大学计算机工程学院) ; College of Science, Mathematics and Technology, Wenzhou-Kean University, Wenzhou, 325060, China(温州-凯恩大学科学、数学与技术学院) ; The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, 511453, China(香港科学与技术大学(广州)) ; School of Professional Studies, New York University, New York, 10003, United States(纽约大学专业研究学院) ; School of Informatics, Xiamen University, Xiamen, 361102, China(厦门大学信息学院)
专题命中 偏好对齐 :RLHF(abstract);prompt injection(abstract);分类 cs.AI
机构 * Dalian University of Technology(大连理工大学) ; University of Surrey(Surrey大学) ; University of Oxford(牛津大学)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG