arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-29 至 2025-07-29 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 4 篇

2507.20150 2025-07-29 cs.AI cs.CL cs.LG 83%

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

Xingcheng Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全训练 :alignment(abstract);RLHF(abstract);safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17275 2025-07-29 eess.SY cs.AI cs.LG cs.SY 81%

Conformal Safety Shielding for Imperfect-Perception Agents

William Scarbro, Calum Imrie, Sinem Getir Yaman, Kavan Fatehi, Corina S. Pasareanu, Radu Calinescu, Ravi Mangal

机构 * Colorado State University, USA(科罗拉多州立大学) University of York, UK(约克大学) Carnegie Mellon University, USA(卡内基梅隆大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 32 pages; Equal contribution by W. Scarbro and C. Imrie; Accepted at 25th International Conference on Runtime Verification, 2025 (RV25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20685 2025-07-29 eess.SY cs.SY 78%

What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems

Marcus Nolte, Nayel Fabian Salem, Olaf Franke, Jan Heckmann, Christoph Höhmann, Georg Stettinger, Markus Maurer

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 1 figure, 1 table, to be published in 2025 IEEE International Automated Vehicle Validation Conference (IAVVC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20614 2025-07-29 cs.CL 57%

Before the Outrage: Challenges and Advances in Predicting Online Antisocial Behavior

Anaïs Ollagnier

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏