arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-26 至 2025-08-26 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 7 篇

2501.18628 2025-08-26 cs.CR cs.AI cs.CL cs.CY 87%

TombRaider: Entering the Vault of History to Jailbreak Large Language Models

Junchen Ding, Jiahao Zhang, Yi Liu, Ziqi Ding, Gelei Deng, Yuekang Li

机构 * UNSW(新南威尔士大学) NTU(国立大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Main Conference of EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11308 2025-08-26 cs.AI cs.CL cs.CR 84%

Defending against Jailbreak through Early Exit Generation of Large Language Models

Chongwen Zhao, Zhihao Dou, Kaizhu Huang

机构 * Duke Kunshan University(杜克昆山大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments ICONIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17710 2025-08-26 cs.CR cs.AI 83%

Optimization-based Prompt Injection Attack to LLM-as-a-Judge

Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, Neil Zhenqiang Gong

机构 * Huazhong University of Science and Technology(华中科技大学) University of Notre Dame(圣母大学) Lehigh University(莱文森大学) Duke University(杜克大学)

专题命中 越狱攻击 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI

Comments To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05945 2025-08-26 cs.CL cs.AI 80%

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

Paul Darm, Annalisa Riccardi

机构 * University of Strathclyde(斯特拉思克莱德大学)

专题命中 越狱攻击 :alignment(abstract,comments);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments Published at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19793 2025-08-26 cs.CR 78%

Prompt Injection Attack to Tool Selection in LLM Agents

Jiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 越狱攻击 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12072 2025-08-26 cs.CR cs.CL 70%

Mitigating Jailbreaks with Intent-Aware LLMs

Wei Jie Yeo, Ranjan Satapathy, Erik Cambria

机构 * Nanyang Technological University(南洋理工大学) Institute of High Performance Computing(高性能计算研究所) Agency for Science, Technology and Research(科技研究局)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17244 2025-08-26 cs.AI 57%

L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems

Aoun E Muhammad, Kin-Choong Yow, Nebojsa Bacanin-Dzakula, Muhammad Attique Khan

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9

详情

展开后加载摘要…

URL PDF HTML 收藏