arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-12 至 2025-09-12 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 2 篇

2405.03486 2025-09-12 cs.CR cs.CV cs.SI 78%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 安全训练 :safety(title,abstract)

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17813 2025-09-12 cs.RO cs.LG 57%

Safe Multi-Agent Navigation guided by Goal-Conditioned Safe Reinforcement Learning

Meng Feng, Viraj Parimi, Brian Williams

机构 * Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(计算机科学与人工智能实验室,麻省理工学院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏