arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-24 至 2025-07-24 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 2 篇

2507.17515 2025-07-24 cs.CV cs.CL 57%

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Songshuo Lu, Hua Wang, Zhi Chen, Yaohua Tang

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05018 2025-07-24 cs.LG cs.MA 57%

Joint Pedestrian and Vehicle Traffic Optimization in Urban Environments using Reinforcement Learning

Bibek Poudel, Xuan Wang, Weizi Li, Lei Zhu, Kevin Heaslip

机构 * Min H. Kao Department of Electrical Engineering and Computer Science at University of Tennessee, Knoxville, TN, USA(田纳西大学电气工程与计算机科学系Min H. Kao部门) Department of Electrical and Computer Engineering at George Mason University(乔治·马歇尔大学电气与计算机工程系) Department of Industrial and Systems Engineering at University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校工业与系统工程系) Department of Civil and Environmental Engineering at University of Tennessee, Knoxville, TN, USA(田纳西大学土木与环境工程系)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏