arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-30 至 2025-07-30 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 3 篇

2507.21820 2025-07-30 cs.CV 85%

Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is

Ahmed B Mustafa, Zihan Ye, Yang Lu, Michael P Pound, Shreyank N Gowda

机构 * School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) Department of Intelligent Science, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学智能科学系) School of Informatics, Xiamen University(厦门大学信息学院)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 70%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21182 2025-07-30 cs.CR cs.AI 70%

SDD: Self-Degraded Defense against Malicious Fine-tuning

Zixuan Chen, Weikai Lu, Xin Lin, Ziqian Zeng

机构 * Zixuan Chen(陈子轩) Weikai Lu(卢伟凯) Xin Lin(林鑫) Ziqian Zeng(曾子谦)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Accepted by ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏