arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1731 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1731 篇

2510.21214 2025-10-27 cs.CR 50%

Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses

Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 越狱攻击 :jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21189 2025-10-27 cs.CR 50%

Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency

Yukun Jiang, Mingjie Li, Michael Backes, Yang Zhang

专题命中 越狱攻击 :jailbreak(abstract)

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13451 2025-10-16 cs.CR 50%

Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts

Li Bai, Qingqing Ye, Xinwei Zhang, Sen Zhang, Zi Liang, Jianliang Xu, Haibo Hu

专题命中 越狱攻击 :alignment(abstract)

Comments To appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11246 2025-10-14 cs.CR 50%

Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems

Pengyu Zhu, Lijun Li, Yaxing Lyu, Li Sun, Sen Su, Jing Shao

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23594 2025-09-30 cs.CR cs.CV 50%

StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data

Yixu Wang, Yan Teng, Yingchun Wang, Xingjun Ma

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 越狱攻击 :safety(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12824 2025-09-18 cs.IR 50%

DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval

Zechao Liu, Zheng Zhou, Xiangkun Chen, Tao Liang, Dapeng Lang

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10563 2025-09-16 cs.CR 50%

Enhancing IoMT Security with Explainable Machine Learning: A Case Study on the CICIOMT2024 Dataset

Mohammed Yacoubi, Omar Moussaoui, C. Drocourt

专题命中 越狱攻击 :safety(abstract)

Journal ref The Third Edition of the International Conference on Connected Objects and Artificial Intelligence (COCIA'2025), Apr 2025, Casablanca (Maroc), Morocco

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04227 2025-09-12 cs.CR 50%

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks

Andreas Happe, Jürgen Cito

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19521 2025-08-22 cs.CR 50%

Security Steerability is All You Need

Itay Hazan, Idan Habler, Ron Bitton, Itsik Mantin

专题命中 越狱攻击 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06154 2025-08-22 cs.CV 50%

GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao, Sivan Doveh, Wei Lin, Paul Gavrikov, Michael Dorkenwald, Shiqi Yang, Saurav Jha, Hiromi Wakaki, Yuki Mitsufuji, Horst Possegger, Rogerio Feris, Leonid Karlinsky, James Glass

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Sony(索尼) Weizmann Institute of Science(魏茨曼科学研究院) JKU(约翰纳斯堡大学) Tübingen AI Center(图宾根人工智能中心) UVA(乌得勒支大学) UNSW(新南威尔士大学) TU Graz(格拉茨技术大学) MIT-IBM(麻省理工学院-IBM)

专题命中 越狱攻击 :safety(abstract)

Comments Code: https://github.com/jmiemirza/GLOV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02332 2025-08-20 cs.CR 50%

PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage

Krishna Kanth Nakka, Xue Jiang, Dmitrii Usynin, Xuebing Zhou

专题命中 越狱攻击 :alignment(abstract)

Comments Preprint. V2 Updated with dataset filtering, benchmarking privacy evaluator and additional latent space visualizations

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13048 2025-08-19 cs.CR 50%

MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies

Weiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng, Puning Zhao, Zhan Qin, Kui Ren

专题命中 越狱攻击 :jailbreak(abstract)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02523 2025-08-05 cs.CR 50%

Transportation Cyber Incident Awareness through Generative AI-Based Incident Analysis and Retrieval-Augmented Question-Answering Systems

Ostonya Thomas, Muhaimin Bin Munir, Jean-Michel Tine, Mizanur Rahman, Yuchen Cai, Khandakar Ashrafi Akbar, Md Nahiyan Uddin, Latifur Khan, Trayce Hockstad, Mashrur Chowdhury

专题命中 越狱攻击 :safety(abstract)

Comments This paper has been submitted to the Transportation Research Board (TRB) for consideration for presentation at the 2026 Annual Meeting

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 越狱攻击 :prompt injection(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19609 2025-07-29 cs.CR 50%

Securing the Internet of Medical Things (IoMT): Real-World Attack Taxonomy and Practical Security Measures

Suman Deb, Emil Lupu, Emm Mic Drakakis, Anil Anthony Bharath, Zhen Kit Leung, Guang Rui Ma, Anupam Chattopadhyay

专题命中 越狱攻击 :safety(abstract)

Comments Submitted as a book chapter in 'Handbook of Industrial Internet of Things' to be published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16576 2025-07-23 cs.CR 50%

From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction

Ahmed Lekssays, Husrev Taha Sencar, Ting Yu

专题命中 越狱攻击 :alignment(abstract)

Comments This paper is accepted at RAID 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12775 2025-07-18 quant-ph 50%

Detecting Entanglement in High-Spin Quantum Systems via a Stacking Ensemble of Machine Learning Models

M. Y. Abd-Rabbou, Amr M. Abdallah, Ahmed A. Zahia, Ashraf A. Gouda, Cong-Feng Qiao

专题命中 越狱攻击 :trustworthy(abstract)

Comments The data and code that support the findings of this study are openly available in a GitHub repository at https://github.com/Amr0MEid/Entanglement-Detection-Using-Ensemble-Learning, and are permanently archived on Zenodo under the DOI:https://doi.org/10.5281/zenodo.16010784

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19697 2025-07-18 cs.CV 50%

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) College of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院)

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10162 2025-07-15 cs.CR 50%

HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning

Weiyang He, Chip-Hong Chang

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13327 2025-07-15 cs.CV 50%

Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning

Ajian Liu, Haocheng Yuan, Xiao Guo, Hui Ma, Wanyi Zhuang, Changtao Miao, Yan Hong, Chuanbiao Song, Jun Lan, Qi Chu, Tao Gong, Yanyan Liang, Weiqiang Wang, Jun Wan, Xiaoming Liu, Zhen Lei

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA)(多模态人工智能系统国家重点实验室(MAIS),自动化研究所,中国科学院(CASIA)) Department of Computer Science, College of Computing, City University of Hong Kong(计算机科学系,计算学院,香港城市大学) School of Computer Science and Engineering, Faculty of Innovation Engineering, Macau University of Science and Technology(计算机科学与工程学院,创新工程学院,澳门科学理工学院) Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学) School of Cyber Science and Technology, University of Science and Technology of China(网络科学与技术学院,中国科学技术大学) Ant Group(蚂蚁集团) Imperial College London(伦敦帝国理工学院)

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20931 2025-06-27 cs.CR 50%

SPA: Towards More Stealth and Persistent Backdoor Attacks in Federated Learning

Chengcheng Zhu, Ye Li, Bosen Rao, Jiale Zhang, Yunlong Mao, Sheng Zhong

专题命中 越狱攻击 :alignment(abstract)

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16699 2025-06-23 cs.CR 50%

Exploring Traffic Simulation and Cybersecurity Strategies Using Large Language Models

Lu Gao, Yongxin Liu, Hongyun Chen, Dahai Liu, Yunpeng Zhang, Jingran Sun

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01518 2025-06-19 cs.CR 50%

Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security

Andrew Adiletta, Zane Weissman, Fatemeh Khojasteh Dana, Berk Sunar, Shahin Tajik

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12411 2025-06-17 cs.CR cs.CV 50%

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

Mengyuan Sun, Yu Li, Yuchen Liu, Bo Du, Yunjie Ge

机构 * 1 School of Cyber Science Engineering, Wuhan University 0.3em 2 School of Computer Science, Wuhan University 0.3em 3 Institute for Math \& AI, Wuhan University

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22429 2025-05-29 cs.CV cs.RO 50%

Zero-Shot 3D Visual Grounding from Vision-Language Models

Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang

机构 * HKUST(GZ)(香港科技大学(广州)) I 2 R, A*STAR(I2R, A*STAR) NUS(国立大学) CSE, HKUST(计算机科学与工程系,香港科技大学)

专题命中 越狱攻击 :alignment(abstract)

Comments 3D-LLM/VLA @ CVPR 2025; Project Page at https://seeground.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23687 2025-05-20 cs.CV cs.CR 50%

Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey

Chiyu Zhang, Lu Zhou, Xiaogang Xu, Jiafei Wu, Zhe Liu

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhejiang Lab, Nanjing University of Aeronautics and Astronautics(浙江实验室,南京航空航天大学)

专题命中 越狱攻击 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08480 2025-04-14 cs.CR 50%

Toward Realistic Adversarial Attacks in IDS: A Novel Feasibility Metric for Transferability

Sabrine Ennaji, Elhadj Benkhelifa, Luigi Vincenzo Mancini

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08205 2025-04-14 cs.CV cs.CR 50%

EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models

Minjae Seo, Myoungsung You, Junhee Lee, Jaehan Kim, Hwanjo Heo, Jintae Oh, Jinwoo Kim

专题命中 越狱攻击 :safety(abstract)

Comments Presented as a poster at ACSAC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21983 2025-04-04 cs.HC cs.SI 50%

Learning to Lie: Reinforcement Learning Attacks Damage Human-AI Teams and Teams of LLMs

Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng, Rasta Tadayon, Xifeng Yan, Ambuj Singh, Francesco Bullo

专题命中 越狱攻击 :safety(abstract)

Comments 17 pages, 9 figures, accepted to ICLR 2025 Workshop on Human-AI Coevolution

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17953 2025-03-25 cs.SE 50%

Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious Prompts

Sheng Ouyang, Yihao Qin, Bo Lin, Liqian Chen, Xiaoguang Mao, Shangwen Wang

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏