arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1731 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1731 篇

2503.08956 2025-03-13 cs.CR 50%

Leaky Batteries: A Novel Set of Side-Channel Attacks on Electric Vehicles

Francesco Marchiori, Mauro Conti

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18549 2025-01-31 cs.CR 50%

CryptoDNA: A Machine Learning Paradigm for DDoS Detection in Healthcare IoT, Inspired by crypto jacking prevention Models

Zag ElSayed, Ahmed Abdelgawad, Nelly Elsayed

专题命中 越狱攻击 :safety(abstract)

Comments 6 pages, 8 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15252 2025-01-28 eess.SP 50%

Deep Multimodal Learning for Real-Time DDoS Attacks Detection in Internet of Vehicles

Mohamed Ababsa, Soheyb Ribouh, Abdelhamid Malki, Lyes Khoukhi

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01508 2025-01-09 physics.soc-ph cs.SI 50%

Garbage in Garbage out: Impacts of data quality on criminal network intervention

Wang Ngai Yeung, Riccardo Di Clemente, Renaud Lambiotte

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11109 2024-12-17 cs.CR 50%

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation

Qinglin Qi, Yun Luo, Yijia Xu, Wenbo Guo, Yong Fang

专题命中 越狱攻击 :jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06255 2024-12-10 cs.CR 50%

Simulation of Multi-Stage Attack and Defense Mechanisms in Smart Grids

Omer Sen, Bozhidar Ivanov, Christian Kloos, Christoph Zol_, Philipp Lutat, Martin Henze, Andreas Ulbig

专题命中 越狱攻击 :safety(abstract)

Journal ref International Journal of Critical Infrastructure Protection 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12513 2024-12-03 cs.CR 50%

Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs

Ahmad Mohsin, Helge Janicke, Adrian Wood, Iqbal H. Sarker, Leandros Maglaras, Naeem Janjua

专题命中 越狱攻击 :safety(abstract)

Comments 27 pages, Standard Journal Paper submitted to Q1 Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13136 2024-11-21 cs.CV 50%

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

Xin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen, Xingjun Ma

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01446 2024-10-31 cs.CV 50%

GuardT2I: Defending Text-to-Image Models from Adversarial Prompts

Yijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong, Qiang Xu

专题命中 越狱攻击 :safety(abstract)

Comments NeurIPS2024 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14005 2024-09-17 cs.CR 50%

Cyber-Twin: Digital Twin-boosted Autonomous Attack Detection for Vehicular Ad-Hoc Networks

Yagmur Yigit, Ioannis Panitsas, Leandros Maglaras, Leandros Tassiulas, Berk Canberk

专题命中 越狱攻击 :safety(abstract)

Comments 6 pages, 5 figures, IEEE International Conference on Communications (ICC) 2024

Journal ref ICC 2024 - IEEE International Conference on Communications, Denver, CO, USA, 2024, pp. 2167-2172

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12076 2024-09-09 cs.CR eess.SP 50%

GAN-GRID: A Novel Generative Attack on Smart Grid Stability Prediction

Emad Efatinasab, Alessandro Brighente, Mirco Rampazzo, Nahal Azadi, Mauro Conti

专题命中 越狱攻击 :safety(abstract)

Journal ref European Symposium on Research in Computer Security (ESORICS 2024), 2024, Lecture Notes in Computer Science, vol 14982

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05264 2024-08-13 cs.CR cs.CV 50%

Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Yihe Fan, Yuxin Cao, Ziyu Zhao, Ziyao Liu, Shaofeng Li

专题命中 越狱攻击 :trustworthy(abstract)

Comments 8 pages, 1 figure. Accepted to 2024 IEEE International Conference on Systems, Man, and Cybernetics

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13111 2024-07-19 cs.MM cs.CV 50%

PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving

Jiyuan Fu, Zhaoyu Chen, Kaixun Jiang, Haijing Guo, Shuyong Gao, Wenqiang Zhang

专题命中 越狱攻击 :safety(abstract)

Comments First-Place in the CVPR 2024 Workshop Challenge: Black-box Adversarial Attacks on Vision Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09292 2024-07-18 cs.CR 50%

Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models

Dong Shu, Mingyu Jin, Tianle Chen, Chong Zhang, Yongfeng Zhang

专题命中 越狱攻击 :safety(abstract)

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06948 2024-07-10 eess.SY cs.SY 50%

Detection-Triggered Recursive Impact Mitigation against Secondary False Data Injection Attacks in Microgrids

Mengxiang Liu, Xin Zhang, Rui Zhang, Zhuoran Zhou, Zhenyong Zhang, Ruilong Deng

专题命中 越狱攻击 :trustworthy(abstract)

Comments Submitted to IEEE Transactions on Smart Grid

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02213 2024-07-09 cs.CV 50%

CosPGD: an efficient white-box adversarial attack for pixel-wise prediction tasks

Shashank Agnihotri, Steffen Jung, Margret Keuper

专题命中 越狱攻击 :alignment(abstract)

Comments Accepted at 41st International Conference on Machine Learning (ICML), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19668 2024-05-31 cs.CV 50%

AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Jiawei Chen, Xiao Yang, Zhengwei Fang, Yu Tian, Yinpeng Dong, Zhaoxia Yin, Hang Su

专题命中 越狱攻击 :jailbreak(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14169 2024-05-24 cs.CV 50%

Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Nhat Chung, Sensen Gao, Tuan-Anh Vu, Jie Zhang, Aishan Liu, Yun Lin, Jin Song Dong, Qing Guo

专题命中 越狱攻击 :safety(abstract)

Comments 12 pages, 5 tables, 5 figures, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12916 2024-04-23 cs.CR 50%

Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models

Zhenyang Ni, Rui Ye, Yuxi Wei, Zhen Xiang, Yanfeng Wang, Siheng Chen

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10299 2024-03-26 cs.CV eess.IV 50%

Boosting Adversarial Transferability by Block Shuffle and Rotation

Kunyu Wang, Xuanran He, Wenxuan Wang, Xiaosen Wang

专题命中 越狱攻击 :trustworthy(abstract)

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12693 2024-03-20 cs.CV 50%

As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?

Anjun Hu, Jindong Gu, Francesco Pinto, Konstantinos Kamnitsas, Philip Torr

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08701 2024-03-20 cs.CR 50%

Review of Generative AI Methods in Cybersecurity

Yagmur Yigit, William J Buchanan, Madjid G Tehrani, Leandros Maglaras

专题命中 越狱攻击 :prompt injection(abstract)

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07654 2024-03-13 cs.IR 50%

Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models

Andrew Parry, Maik Fröbe, Sean MacAvaney, Martin Potthast, Matthias Hagen

专题命中 越狱攻击 :prompt injection(abstract)

Comments 13 pages, 3 figures, Accepted at ECIR 2024 as a Full Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15617 2024-02-27 cs.CR cs.SY eess.SY 50%

Reinforcement Learning-Based Approaches for Enhancing Security and Resilience in Smart Control: A Survey on Attack and Defense Methods

Zheyu Zhang

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00994 2024-01-03 cs.CR 50%

Detection and Defense Against Prominent Attacks on Preconditioned LLM-Integrated Virtual Assistants

Chun Fai Chan, Daniel Wankit Yip, Aysan Esmradi

专题命中 越狱攻击 :safety(abstract)

Comments Accepted to be published in the Proceedings of the 10th IEEE CSDE 2023, the Asia-Pacific Conference on Computer Science and Data Engineering 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15736 2023-12-29 cs.CR cs.SY eess.SY 50%

Vulnerability of Machine Learning Approaches Applied in IoT-based Smart Grid: A Review

Zhenyong Zhang, Mengxiang Liu, Mingyang Sun, Ruilong Deng, Peng Cheng, Dusit Niyato, Mo-Yuen Chow, Jiming Chen

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04403 2023-12-08 cs.CV 50%

OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization

Dongchen Han, Xiaojun Jia, Yang Bai, Jindong Gu, Yang Liu, Xiaochun Cao

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18350 2023-12-01 cs.DC 50%

Unveiling Backdoor Risks Brought by Foundation Models in Heterogeneous Federated Learning

Xi Li, Chen Wu, Jiaqi Wang

专题命中 越狱攻击 :safety(abstract)

Comments Jiaqi Wang is the corresponding author. arXiv admin note: text overlap with arXiv:2311.00144

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12685 2023-11-22 cs.CV 50%

Rethinking the Backward Propagation for Adversarial Transferability

Xiaosen Wang, Kangheng Tong, Kun He

专题命中 越狱攻击 :trustworthy(abstract)

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13345 2023-10-23 cs.CR 50%

An LLM can Fool Itself: A Prompt-Based Adversarial Attack

Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, Mohan Kankanhalli

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏