arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-27 至 2026-01-27 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 4 篇

2601.17549 2026-01-27 cs.CR cs.AI 79%

Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents

打破协议:模型上下文协议规范的安全分析及工具集成LLM代理中的提示注入漏洞

Narek Maloyan, Dmitry Namiot

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文分析了模型上下文协议的安全漏洞,提出MCPSec协议扩展以提升安全性,减少攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17383 2026-01-27 cs.CV cs.AI 79%

Physical Prompt Injection Attacks on Large Vision-Language Models

针对大视觉-语言模型的物理提示注入攻击

Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本研究提出了一种无需访问模型或输入的物理提示注入攻击方法,通过物理物体注入恶意指令,成功攻击多种大视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14005 2026-01-27 cs.CR cs.LG 79%

PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features

PIShield: 通过内在LLM特性检测提示注入攻击

Wei Zou, Yupei Liu, Yanting Wang, Ying Chen, Neil Gong, Jinyuan Jia

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

AI总结 PIShield通过利用指令微调LLM的内部表示,高效检测提示注入攻击,显著优于现有方法。

Comments The code is available at https://github.com/weizou52/PIShield

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17548 2026-01-27 cs.CR 78%

Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems

针对代理编码助手的提示注入攻击:对技能、工具和协议生态系统漏洞的系统分析

Narek Maloyan, Dmitry Namiot

专题命中 提示注入 :prompt injection(title,abstract)

AI总结 本文系统分析了代理编码助手的提示注入攻击漏洞,提出三维分类法并揭示防御不足,强调需架构级防护而非碎片化措施。

详情

展开后加载摘要…

URL PDF HTML 收藏