PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
PVDetector:通过策略违规概念分析检测针对特定目的大语言模型代理的提示注入攻击
AI总结 研究针对特定目的LLM代理的提示注入攻击,提出PVDetector框架,通过测量与离线派生的PV概念的隐藏状态对齐来检测攻击,实验表明该方法误报率低、开销小,性能优于现有方法。
Comments Accepted to ACM MM 2026. Code: https://github.com/Claresigle/PVDetector