arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-21 至 2025-10-21 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 6 篇

2510.15430 2025-10-21 cs.CV cs.AI 85%

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17006 2025-10-21 cs.CL 83%

Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization

Masahiro Kaneko, Zeerak Talat, Timothy Baldwin

机构 * MBZUAI(马克斯·普朗克人工智能研究所) University of Edinburgh(爱丁堡大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17000 2025-10-21 cs.CR cs.CL cs.LG 73%

Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs

Masahiro Kaneko, Timothy Baldwin

机构 * MBZUAI Abu Dhabi, UAE(阿布扎赫德MBZUAI)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19360 2025-10-21 cs.CL cs.AI 62%

Semantic Representation Attack against Aligned Large Language Models

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系) School of Electronics and Information, Northwestern Polytechnical University(西北工业大学电子与信息学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19304 2025-10-21 cs.CY cs.AI cs.CE 62%

Epistemic Trade-Off: An Analysis of the Operational Breakdown and Ontological Limits of "Certainty-Scope" in AI

Generoso Immediato

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY

Comments Preprint V3 (October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16794 2025-10-21 cs.CR cs.LG 57%

Black-box Optimization of LLM Outputs by Asking for Directions

Jie Zhang, Meng Ding, Yang Liu, Jue Hong, Florian Tramèr

机构 * ETH Zurich(苏黎世联邦理工学院) University at Buffalo(布法罗大学) Bytedance, Security Research(字节跳动安全研究)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏