arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1731 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1731 篇

2505.13899 2025-09-30 cs.LG 57%

Causes and Consequences of Representational Similarity in Machine Learning Models

Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo, Emily Wenger

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17371 2025-09-24 cs.CR cs.LG 57%

SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models

Haotian Xu, Qingsong Peng, Jie Shi, Huadi Zheng, Yu Li, Cheng Zhuo

机构 * Zhejiang University(浙江大学) Huawei(华为)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15739 2025-09-22 cs.CL 57%

Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics

Reza Sanayei, Srdjan Vesic, Eduardo Blanco, Mihai Surdeanu

机构 * Department of Computer Science, University of Arizona(亚利桑那大学计算机科学系) CRIL CNRS & University of Artois(CNRS CRIL与阿维尼昂大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02844 2025-09-17 cs.CV cs.CL cs.CR 57%

Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection

Ziqi Miao, Yi Ding, Lijun Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main). 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15244 2025-09-17 cs.CV cs.AI 57%

Adversarial Prompt Distillation for Vision-Language Models

Lin Luo, Xin Wang, Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07941 2025-09-10 cs.CR cs.AI 57%

ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation

Kai Ye, Liangcai Su, Chenxiong Qian

机构 * The University of Hong Kong(香港大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

Comments This paper has been accepted by the ACM Conference on Computer and Communications Security (CCS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19288 2025-08-28 cs.CR cs.AI 57%

Tricking LLM-Based NPCs into Spilling Secrets

Kyohei Shiomi, Zhuotao Lian, Toru Nakanishi, Teruaki Kitasuka

机构 * Hiroshima University(广岛大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17244 2025-08-26 cs.AI 57%

L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems

Aoun E Muhammad, Kin-Choong Yow, Nebojsa Bacanin-Dzakula, Muhammad Attique Khan

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15283 2025-08-22 cs.IR cs.CL 57%

Adversarial Attacks against Neural Ranking Models via In-Context Learning

Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, Charles L. A. Clarke

机构 * University of Waterloo(多伦多大学) University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20841 2025-08-19 cs.CL 57%

Concealment of Intent: A Game-Theoretic Analysis

Xinbo Wu, Abhishek Umrawal, Lav R. Varshney

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02141 2025-08-13 cs.CV cs.CL 57%

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, Linlin Shen

机构 * Shenzhen University(深圳大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) City University of Hong Kong(香港城市大学) Stanford University(斯坦福大学) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments ICCV 2025, 38 pages, 22 figures, 35 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04632 2025-08-08 cs.CL 57%

IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards

Xu Guo, Tianyi Liang, Tong Jian, Xiaogui Yang, Ling-I Wu, Chenhui Li, Zhihui Lu, Qipeng Guo, Kai Chen

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03125 2025-08-06 cs.CR cs.AI cs.MA 57%

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS

Bingyu Yan, Ziyi Zhou, Xiaoming Zhang, Chaozhuo Li, Ruilin Zeng, Yirui Qi, Tianbo Wang, Litian Zhang

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03110 2025-08-06 cs.CL 57%

Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation

Zizhong Li, Haopeng Zhang, Jiawei Zhang

机构 * University of California, Davis(加州大学戴维斯分校) University of Hawaii at Mānoa(夏威夷大学马诺阿分校)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19880 2025-07-29 cs.CR cs.AI 57%

Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data

Nicola Croce, Tobin South

机构 * Pivotal Research(Pivotal研究机构) Stanford University(斯坦福大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Abstract submitted to the Technical AI Governance Forum 2025 (https://www.techgov.ai/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18656 2025-07-28 cs.CV cs.LG 57%

ShrinkBox: Backdoor Attack on Object Detection to Disrupt Collision Avoidance in Machine Learning-based Advanced Driver Assistance Systems

Muhammad Zaeem Shahzad, Muhammad Abdullah Hanif, Bassem Ouni, Muhammad Shafique

机构 * eBRAIN Lab, New York University Abu Dhabi (NYUAD), UAE(eBRAIN实验室,纽约大学阿布扎赫尔分校(NYUAD),阿联酋) AI and Digital Science Research Center, Technology Innovation Institute (TII), Abu Dhabi, UAE(人工智能与数字科学研究中心,技术创新研究所(TII),阿布扎赫尔,阿联酋)

专题命中 越狱攻击 :safety(abstract);分类 cs.LG

Comments 8 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13761 2025-07-21 cs.CL 57%

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

Palash Nandi, Maithili Joshi, Tanmoy Chakraborty

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11117 2025-07-16 cs.AI 57%

AI Agent Architecture for Decentralized Trading of Alternative Assets

Ailiya Borjigin, Cong He, Charles CC Lee, Wei Zhou

机构 * Centre for Sustainable Development, University of Newcastle (Australia), Singapore(可持续发展中心,新南威尔士大学(澳大利亚),新加坡)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments 8 Pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06323 2025-07-10 cs.CR cs.AI 57%

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms

Tarek Gasmi, Ramzi Guesmi, Ines Belhadj, Jihene Bennaceur

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10273 2025-07-09 cs.CR cs.AI cs.NI 57%

AttentionGuard: Transformer-based Misbehavior Detection for Secure Vehicular Platoons

Hexu Li, Konstantinos Kalogiannis, Ahmed Mohamed Hussain, Panos Papadimitratos

机构 * Networked Systems Security (NSS) Group\ Royal Institute of Technology Stockholm Sweden Networked Systems Security (NSS) Group\ Royal Institute of Technology

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments Author's version; Accepted for presentation at the ACM Workshop on Wireless Security and Machine Learning (WiseML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02057 2025-07-04 cs.CR cs.AI 57%

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

Lu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An, Guangyu Shen, Zhou Xuan, Xuan Chen, Xiangyu Zhang

机构 * Purdue University(普渡大学) Columbia University(哥伦比亚大学) Virginia Tech(弗吉尼亚理工大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00052 2025-07-02 cs.CV cs.AI 57%

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven, West Haven, CT, USA(SAIL实验室,新罕布什尔大学,西哈文,康涅狄格州,美国)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21842 2025-06-30 quant-ph cs.CR cs.LG 57%

Adversarial Threats in Quantum Machine Learning: A Survey of Attacks and Defenses

Archisman Ghosh, Satwik Kundu, Swaroop Ghosh

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.LG

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17318 2025-06-24 cs.CR cs.AI 57%

Context manipulation attacks : Web agents are susceptible to corrupted memory

Atharv Singh Patlan, Ashwin Hebbar, Pramod Viswanath, Prateek Mittal

机构 * princeton(普林斯顿大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11934 2025-06-17 cs.AI 57%

Stepwise Reasoning Error Disruption Attack of LLMs

Jingyu Peng, Maolin Wang, Xiangyu Zhao, Kai Zhang, Wanyu Wang, Pengyue Jia, Qidong Liu, Ruocheng Guo, Qi Liu

机构 * University of Science and Technology of China(中国科学技术大学) City University of Hong Kong(香港城市大学) Independent Researcher(独立研究者)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11521 2025-06-16 cs.CR cs.AI cs.MM 57%

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

Jinming Wen, Xinyi Wu, Shuai Zhao, Yanhao Jia, Yuwen Li

机构 * Jilin University(吉林大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Northeastern University(东北大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07948 2025-06-10 cs.LG cs.CR 57%

TokenBreak: Bypassing Text Classification Models Through Token Manipulation

Kasimir Schulz, Kenneth Yeung, Kieran Evans

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07987 2025-06-06 cs.AI 57%

Universal Adversarial Attack on Aligned Multimodal LLMs

Temurbek Rahmatullaev, Polina Druzhinina, Nikita Kurdiukov, Matvey Mikhalchuk, Andrey Kuznetsov, Anton Razzhigaev

机构 * AIRI MSU(莫斯科国立大学) HSE University(俄罗斯高等经济大学) Skoltech(斯克里普钦科技大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

Comments Added benchmarks, baselines, author, appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02859 2025-06-04 cs.CR cs.AI 57%

ATAG: AI-Agent Application Threat Assessment with Attack Graphs

Parth Atulbhai Gandhi, Akansha Shukla, David Tayouri, Beni Ifland, Yuval Elovici, Rami Puzis, Asaf Shabtai

机构 * Dept. of Software and Information Systems Engineering(软件与信息系统工程系) Ben-Gurion University of the Negev(贝叶尔-加利利大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13715 2025-06-03 cs.CR cs.CY cs.HC 57%

Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing

Marc Schmitt, Ivan Flechais

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CY

Comments Submitted to CHI 2024

Journal ref Artificial Intelligence Review, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏