arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1731 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1731 篇

2504.01094 2025-04-03 cs.SD cs.AI cs.CL cs.CR eess.AS 62%

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

Jaechul Roh, Virat Shejwalkar, Amir Houmansadr

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21464 2025-03-28 cs.CL cs.AI cs.PF 62%

Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection

Ryan Marinelli, Josef Pichlmeier, Tamas Bisztray

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06608 2025-03-25 cs.LG cs.AI cs.CR 62%

MF-CLIP: Leveraging CLIP as Surrogate Models for No-box Adversarial Attacks

Jiaming Zhang, Lingyu Qiu, Qi Yi, Yige Li, Jitao Sang, Changsheng Xu, Dit-Yan Yeung

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08226 2025-03-12 cs.CL cs.AI 62%

A Grey-box Text Attack Framework using Explainable AI

Esther Chiramal, Kelvin Soh Boon Kai

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19038 2025-03-11 cs.CL cs.LG 62%

DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs

Ben Ganon, Alon Zolfi, Omer Hofman, Inderjeet Singh, Hisashi Kojima, Yuval Elovici, Asaf Shabtai

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14628 2025-02-21 cs.LG cs.CL 62%

PEARL: Towards Permutation-Resilient LLMs

Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09039 2025-01-17 cs.CR cs.AI cs.CY 62%

Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models

Abdulkadir Erol, Trilok Padhi, Agnik Saha, Ugur Kursuncu, Mehmet Emin Aktas

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14795 2025-01-09 cs.CL cs.CR cs.LG 62%

Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models

Jiaming He, Wenbo Jiang, Guanyu Hou, Wenshu Fan, Rui Zhang, Hongwei Li

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments The paper has been accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11208 2024-10-30 cs.CR cs.AI cs.CL 62%

Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, Xu Sun

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2024, camera ready version. Code and data are available at https://github.com/lancopku/agent-backdoor-attacks

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06360 2024-10-28 cs.LG cs.AI 62%

Exploring the Landscape of Machine Unlearning: A Comprehensive Survey and Taxonomy

Thanveer Shaik, Xiaohui Tao, Haoran Xie, Lin Li, Xiaofeng Zhu, Qing Li

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14774 2024-10-24 cs.CV cs.CL cs.CR cs.LG 62%

Few-Shot Adversarial Prompt Learning on Vision-Language Models

Yiwei Zhou, Xiaobo Xia, Zhiwei Lin, Bo Han, Tongliang Liu

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07242 2024-10-22 cs.CR cs.AI cs.CL 62%

Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs

Bibek Upadhayay, Vahid Behzadan

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13968 2024-10-11 cs.CL cs.AI cs.CR 62%

Protecting Your LLMs with Information Bottleneck

Zichuan Liu, Zefan Wang, Linjie Xu, Jinyu Wang, Lei Song, Tianchun Wang, Chunlin Chen, Wei Cheng, Jiang Bian

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments Accepted by Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08755 2024-10-10 cs.CR cs.AI cs.LG 62%

Distributed Threat Intelligence at the Edge Devices: A Large Language Model-Driven Approach

Syed Mhamudul Hasan, Alaa M. Alotaibi, Sajedul Talukder, Abdur R. Shahid

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05949 2024-10-10 cs.CL cs.AI cs.CR 62%

Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning

Shuai Zhao, Meihuizi Jia, Luu Anh Tuan, Fengjun Pan, Jinming Wen

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14866 2024-10-08 cs.CL cs.CR cs.LG 62%

Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models

Hongfu Liu, Yuxi Xie, Ye Wang, Michael Shieh

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted to the EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19173 2024-10-01 cs.CL cs.AI 62%

HM3: Heterogeneous Multi-Class Model Merging

Stefan Hackmann

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04822 2024-09-10 cs.CL cs.AI 62%

Exploring Straightforward Conversational Red-Teaming

George Kour, Naama Zwerdling, Marcel Zalmanovici, Ateret Anaby-Tavor, Ora Nova Fandina, Eitan Farchi

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17235 2024-09-06 cs.CR cs.AI cs.LG 62%

AI-Driven Intrusion Detection Systems (IDS) on the ROAD Dataset: A Comparative Analysis for Automotive Controller Area Network (CAN)

Lorenzo Guerra, Linhan Xu, Paolo Bellavista, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10107 2024-08-20 cs.LG cs.AI stat.ML 62%

Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments

Heeyoung Lee, Hoyoon Byun, Changdae Oh, JinYeong Bak, Kyungwoo Song

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to European Conference on Artificial Intelligence (ECAI) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15399 2024-07-23 cs.CL cs.AI cs.CR 62%

Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models

Xiao Liu, Liangzhi Li, Tong Xiang, Fuying Ye, Lu Wei, Wangyue Li, Noa Garcia

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13796 2024-07-22 cs.CR cs.AI cs.CL 62%

Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models

Zihao Xu, Yi Liu, Gelei Deng, Kailong Wang, Yuekang Li, Ling Shi, Stjepan Picek

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12822 2024-07-19 cs.CL cs.AI 62%

Lightweight Large Language Model for Medication Enquiry: Med-Pal

Kabilan Elangovan, Jasmine Chiat Ling Ong, Liyuan Jin, Benjamin Jun Jie Seng, Yu Heng Kwan, Lit Soo Tan, Ryan Jian Zhong, Justina Koi Li Ma, YuHe Ke, Nan Liu, Kathleen M Giacomini, Daniel Shu Wei Ting

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03160 2024-07-04 cs.CR cs.CL cs.LG 62%

SOS! Soft Prompt Attack Against Open-Source Large Language Models

Ziqing Yang, Michael Backes, Yang Zhang, Ahmed Salem

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15518 2024-06-25 cs.CL cs.LG 62%

Steering Without Side Effects: Improving Post-Deployment Control of Language Models

Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau, Salsabila Mahdi, Samuel R. Bowman

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14857 2024-06-21 cs.CL cs.AI cs.CR 62%

Is the System Message Really Important to Jailbreaks in Large Language Models?

Xiaotian Zou, Yongkang Chen, Ke Li

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 13 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00240 2024-06-04 cs.LG cs.CL cs.CR 62%

Exploring Vulnerabilities and Protections in Large Language Models: A Survey

Frank Weizhen Liu, Chenhui Hu

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15245 2024-05-27 cs.LG cs.AI 62%

Cooperative Backdoor Attack in Decentralized Reinforcement Learning with Theoretical Guarantee

Mengtong Gao, Yifei Zou, Zuyuan Zhang, Xiuzhen Cheng, Dongxiao Yu

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07308 2024-05-03 cs.CL cs.AI 62%

LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, Duen Horng Chau

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02406 2024-04-04 cs.CR cs.AI cs.CL 62%

Exploring Backdoor Vulnerabilities of Chat Models

Yunzhuo Hao, Wenkai Yang, Yankai Lin

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

Comments Code and data are available at https://github.com/hychaochao/Chat-Models-Backdoor-Attacking

详情

展开后加载摘要…

URL PDF HTML 收藏