arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-10 至 2025-11-10 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 3 篇

2504.21700 2025-11-10 cs.CR cs.AI cs.LG 81%

XBreaking: Understanding how LLMs security alignment can be broken

Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Vinod P

机构 * Department of Electrical, Computer Biomedical Engineering, University of Pavia, Italy\ . Department of Computer Applications,\ University of Science \& Technology, India\ .

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18469 2025-11-10 cs.CL cs.LG 79%

Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities

Chung-En Sun, Xiaodong Liu, Weiwei Yang, Tsui-Wei Weng, Hao Cheng, Aidan San, Michel Galley, Jianfeng Gao

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments Accepted to NAACL 2025 Main (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03299 2025-11-10 cs.LG cs.CL cs.CV 62%

GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models

Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang

机构 * School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校) Independent Researcher, Starc Institute(Starc研究所独立研究者) Computer Science and Engineering HKUST(HKUST计算机科学与工程学院) Computer Science Lapis Labs University of Illinois Urbana-Champaign(计算机科学Lapis Labs伊利诺伊大学厄巴纳-香槟分校) School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments 28 papges

详情

展开后加载摘要…

URL PDF HTML 收藏