URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
专题命中 安全训练 :alignment(abstract);分类 cs.CL
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全训练 :alignment(abstract);分类 cs.CL
机构 * Min H. Kao Department of Electrical Engineering and Computer Science at University of Tennessee, Knoxville, TN, USA(田纳西大学电气工程与计算机科学系Min H. Kao部门) ; Department of Electrical and Computer Engineering at George Mason University(乔治·马歇尔大学电气与计算机工程系) ; Department of Industrial and Systems Engineering at University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校工业与系统工程系) ; Department of Civil and Environmental Engineering at University of Tennessee, Knoxville, TN, USA(田纳西大学土木与环境工程系)
专题命中 安全训练 :safety(abstract);分类 cs.LG
Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025