arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PVDetector:通过策略违规概念分析检测针对特定目的大语言模型代理的提示注入攻击

PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis

Junhui Wang, Hangtao Zhang, Zhirun Zheng, Li Zeng, Jiejun Xiao, Xi Luo, Lihua Yin, Saiqin Long

arXiv 2607.12624首次发表:更新:

AI 中文总结

研究针对特定目的LLM代理的提示注入攻击,提出PVDetector框架,通过测量与离线派生的PV概念的隐藏状态对齐来检测攻击,实验表明该方法误报率低、开销小,性能优于现有方法。

AI 中文摘要

大语言模型(LLMs)越来越多地被部署为特定目的代理来处理客户服务和代码生成等领域特定任务。这些代理不仅要遵守通用安全护栏,还要遵循特定于其指定角色的限制,这扩大了攻击面,尤其是提示注入(PI)攻击。现有检测方法主要依赖分析输入输出模式,效果有限。我们通过分析隐藏激活空间发现,当用超出其指定目的的请求提示时,LLMs固有地保留潜在策略违规(PV)概念。基于此,我们提出PVDetector,一个无训练框架,通过测量与PV概念的隐藏状态对齐来检测LLM推理期间的PI攻击,PV概念从违规和合规提示的对比对离线派生。跨多个LLMs和数据集的实验表明,PVDetector实现了<1%的误报率,辅助开销最小,始终优于现有方法。

英文摘要

Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but also purpose-specific restrictions tailored to their designated roles. Such additional restrictions enlarge the attack surface, particularly to prompt injection (PI) attacks. To defend against such attacks, existing detection methods primarily rely on analyzing input-output patterns, yet yield limited effectiveness. To address this limitation, we turn to analyzing the hidden activation space and discover that LLMs inherently retain latent policy-violation (PV) concepts when prompted with requests beyond their designated purpose. Particularly, PV concepts capture the semantics of conflicts between user queries and predefined restrictions, implicitly reflecting LLMs' intrinsic awareness of recognizing policy violations. Building on this insight, we propose PVDetector, a training-free framework that detects PI attacks during LLM inference by measuring hidden-state alignment with PV concepts, which are derived offline from the contrastive pairs of policy-violating and policy-compliant prompts. Experiments across multiple LLMs and datasets show that PVDetector achieves <1\% false negative rate with minimal auxiliary overhead, consistently outperforming state-of-the-art methods. Our code is available at https://github.com/Claresigle/PVDetector .

CommentsAccepted to ACM MM 2026. Code: https://github.com/Claresigle/PVDetector

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑