arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-22 至 2025-10-22 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 3 篇

2510.18728 2025-10-22 cs.CR cs.AI 79%

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

Sidhant Narula, Javad Rafiei Asl, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi

机构 * University of Arizona(亚利桑那大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments This paper has been accepted for presentation at the Conference on Applied Machine Learning in Information Security (CAMLIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03417 2025-10-22 cs.CR cs.AI 70%

NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks

Javad Rafiei Asl, Sidhant Narula, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi

机构 * Old Dominion University(旧 Dominion 大学) University of Arizona(亚利桑那大学)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI

Comments This paper has been accepted in the main conference proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025). Javad Rafiei Asl and Sidhant Narula are co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07153 2025-10-22 cs.CR cs.AI 57%

Mind the Web: The Security of Web Use Agents

Avishag Shapira, Parth Atulbhai Gandhi, Edan Habler, Asaf Shabtai

机构 * Ben-Gurion University of the Negev, Israel(内盖夫本·古里安大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏