arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-18 至 2025-12-18 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 2 篇

2507.07417 2025-12-18 cs.CR cs.AI cs.CL 81%

May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks

请允许我注意:利用架构感知攻击打破基于微调的提示注入防御

Nishit V. Pandya, Andrey Labunets, Sicun Gao, Earlence Fernandes

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种基于注意力的攻击算法,揭示了现有白盒提示注入防御的不足,攻击成功率高达85-95%

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14285 2025-12-18 cs.CR cs.LG 79%

A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks

针对提示注入攻击的多智能体LLM防御管道

S M Asif Hossain, Ruksat Khan Shayoni, Mohd Ruhul Ameen, Akif Islam, M. F. Mridha, Jungpil Shin

机构 * School of Computing, Wichita State University, Kansas, USA(威斯康星州立大学计算机学院) College of Engineering and Computer Sciences, Marshall University, Huntington, WV, USA(马歇尔大学工程与计算机科学学院) Department of Computer Science and Engineering, University of Rajshahi, Bangladesh(拉贾沙希大学计算机科学与工程系) Department of Computer Science and Engineering, American International University-Bangladesh, Dhaka, Bangladesh(美国国际大学-孟加拉国计算机科学与工程系) School of Computer Science and Engineering, The University of Aizu, Aizuwakamatsu, Japan(立命馆大学计算机科学与工程学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

AI总结 本文提出了一种多智能体防御框架,通过协调的LLM代理实时检测并中和提示注入攻击,显著提升了安全性和系统功能。

Comments Accepted at the 11th IEEE WIECON-ECE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏