arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1731 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1731 篇

2509.19100 2025-09-24 cs.LG cs.AI 62%

Algorithms for Adversarially Robust Deep Learning

Alexander Robey

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 62%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16792 2025-09-23 cs.CL cs.AI 62%

MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning

Muyang Zheng, Yuanzhi Yao, Changting Lin, Caihong Kai, Yanxiang Chen, Zhiquan Liu

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Cyber Security, Jinan University(暨南大学网络安全学院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13266 2025-09-17 cs.LG cs.AI 62%

JANUS: A Dual-Constraint Generative Framework for Stealthy Node Injection Attacks

Jiahao Zhang, Xiaobing Pei, Zhaokun Zhong, Wenqiang Hao, Zhenghao Tang

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10931 2025-09-16 cs.AI cs.CL 62%

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

Seongho Joo, Hyukhun Koh, Kyomin Jung

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17674 2025-09-10 cs.CR cs.AI cs.LG 62%

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

Qiming Guo, Jinwen Tang, Xingran Huang

机构 * Department of Computer Science Texas A\&M University–Corpus Christi Corpus Christi, TX, USA EECS Department University of Missouri Columbia, MO, USA Department of Computer Engineering University of California–Riverside Riverside, CA, USA

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10722 2025-09-03 cs.CL cs.AI 62%

MEGen: Generative Backdoor into Large Language Models via Model Editing

Jiyang Qiu, Xinbei Ma, Zhuosheng Zhang, Hai Zhao, Yun Li, Qianren Wang

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering, Shanghai Jiao Tong University(上海交通大学智能交互与认知工程重点实验室) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(上海Web3可信数据流通与治理重点实验室) Cognitive AI Lab(认知人工智能实验室)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11727 2025-09-03 cs.CR cs.AI cs.CL cs.SE 62%

Efficient Detection of Toxic Prompts in Large Language Models

Yi Liu, Junzhe Yu, Huijia Sun, Ling Shi, Gelei Deng, Yuqi Chen, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) ShanghaiTech University(上海科技大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted by the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04196 2025-08-07 cs.CL cs.AI cs.CR 62%

Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models

Siddhant Panpatil, Hiskias Dingeto, Haon Park

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17922 2025-07-25 cs.LG cs.AI 62%

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models

Jessica Quaye, Charvi Rastogi, Alicia Parrish, Oana Inel, Minsuk Kahng, Lora Aroyo, Vijay Janapa Reddi

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17946 2025-07-10 cs.CR cs.AI cs.CL 62%

Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs

Shuai Zhao, Leilei Gan, Zhongliang Guo, Xiaobao Wu, Yanhao Jia, Luwei Xiao, Cong-Duy Nguyen, Luu Anh Tuan

机构 * Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) University of St Andrews(圣安德鲁大学) East China Normal University(华东师范大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00406 2025-07-09 cs.AI cs.CL 62%

Agents Are All You Need for LLM Unlearning

Debdeep Sanyal, Murari Mandal

机构 * RespAI Lab, School of Computer Engineering, KIIT Bhubaneswar(RespAI实验室,计算机工程学院,KIIT大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16327 2025-07-09 cs.CR cs.AI cs.CL 62%

Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs

Rui Pu, Chaozhuo Li, Rui Ha, Zejian Chen, Litian Zhang, Zheng Liu, Lirong Qiu, Zaisheng Ye

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Hangzhou Dianzi University(杭州电子科技大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Fujian Cancer Hospital(福建癌症医院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00239 2025-07-02 cs.CL cs.AI 62%

Linearly Decoding Refused Knowledge in Aligned Language Models

Aryan Shrivastava, Ari Holtzman

机构 * University of Chicago(芝加哥大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04202 2025-06-27 cs.CR cs.AI cs.LG 62%

TracLLM: A Generic Framework for Attributing Long Context LLMs

Yanting Wang, Wei Zou, Runpeng Geng, Jinyuan Jia

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI、cs.LG

Comments To appear in USENIX Security Symposium 2025. The code and data are at: https://github.com/Wang-Yanting/TracLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01633 2025-06-26 cs.LG cs.AI 62%

Adversarial Reasoning at Jailbreaking Time

Mahdi Sabbaghi, Paul Kassianik, George Pappas, Yaron Singer, Amin Karbasi, Hamed Hassani

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22960 2025-06-23 cs.AI cs.LG 62%

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Yongjin Yang, Euiin Yi, Jongwoo Ko, Kimin Lee, Zhijing Jin, Se-Young Yun

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13726 2025-06-17 cs.AI cs.CR cs.LG 62%

Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

Arjun Krishna, Aaditya Rastogi, Erick Galinkin

机构 * University of Waterloo(滑铁卢大学) NVIDIA(NVIDIA公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to LLMSEC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00382 2025-06-06 cs.LG cs.CL 62%

Spectral Insights into Data-Oblivious Critical Layers in Large Language Models

Xuyuan Liu, Lei Hsiung, Yaoqing Yang, Yujun Yan

机构 * Dartmouth College(达特茅斯学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by Findings of ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01926 2025-06-04 cs.CL cs.AI 62%

Unnatural Languages Are Not Bugs but Features for LLMs

Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu, Tianle Cai, Longxu Dou, Kenji Kawaguchi, Anirudh Goyal, J. Zico Kolter, Michael Qizhe Shieh

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00548 2025-06-03 cs.CR cs.CL cs.LG 62%

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities

Jiahui Geng, Thy Thy Tran, Preslav Nakov, Iryna Gurevych

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05962 2025-06-03 cs.CL cs.AI 62%

Effective faking of verbal deception detection with target-aligned adversarial attacks

Bennett Kleinberg, Riccardo Loconte, Bruno Verschuere

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to Legal and Criminological Psychology (author version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24232 2025-06-02 cs.CV cs.AI cs.CL 62%

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models

Haibo Jin, Peiyan Zhang, Peiran Wang, Man Luo, Haohan Wang

机构 * School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校) Computer Science and Engineering HKUST(计算机科学与工程香港科技大学) Computer Science Department University of California, Los Angeles(计算机科学系加州大学洛杉矶分校) Research Scientist, Intel Labs(英特尔实验室研究员) School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17541 2025-05-30 cs.AI cs.CL 62%

Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction

Michal Bravansky, Vaclav Kubon, Suhas Hariharan, Robert Kirk

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16359 2025-05-30 cs.CL cs.AI 62%

Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context

Nilanjana Das, Edward Raff, Aman Chadha, Manas Gaur

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Amazon Web Services(亚马逊网络服务)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2407.14644

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10619 2025-05-29 cs.AI cs.CL cs.CR 62%

Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search

Andy Zhou, Ron Arel

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17040 2025-05-26 cs.LG cs.CL 62%

Generalizing Large Language Model Usability Across Resource-Constrained

Yun-Da Tsai

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG

Comments Doctoral disstertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14425 2025-05-21 cs.CL cs.AI cs.CR 62%

Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation

Shuai Zhao, Xiaobao Wu, Cong-Duy Nguyen, Yanhao Jia, Meihuizi Jia, Yichao Feng, Luu Anh Tuan

机构 * Nanyang Technological University(南洋理工大学) Northwest Normal University(西北师范大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04578 2025-05-08 cs.LG cs.AI 62%

Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization

Wenjun Cao

机构 * Independent Researcher(独立研究者)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00108 2025-05-02 cs.CR cs.AI cs.CL 62%

LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem

Hongyi Liu, Shaochen Zhong, Xintong Sun, Minghao Tian, Mohsen Hariri, Zirui Liu, Ruixiang Tang, Zhimeng Jiang, Jiayi Yuan, Yu-Neng Chuang, Li Li, Soo-Hyun Choi, Rui Chen, Vipin Chaudhary, Xia Hu

机构 * Rice University(里士大学) Case Western Reserve University(凯斯西储大学) University of Minnesota(明尼苏达大学) Rutgers University(罗格斯大学) Texas A&M University(德克萨斯阿姆大学) Samsung Electronics America(三星电子美国公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏