arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30657cs.CRcs.CL

通过攻击链建模检测电子邮件智能体中的提示注入

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

Ahmad Hashmi, Dhyey Patel, Yunting Yin

首次发表
浏览论文内容

中文总结 AI 辅助

针对电子邮件智能体的间接提示注入,提出基于攻击链建模的检测框架,结合多阶段验证与逻辑策略,在五个基准上显著优于现有检测器。

中文摘要 AI 辅助

大语言模型电子邮件助手特别容易受到间接提示注入的攻击,因为不受信任的电子邮件内容可能被检索到模型上下文中,并影响后续的工具使用。现有的提示注入检测器主要将这一问题表述为二元恶意文本分类,这忽略了有害智能体行为往往通过一系列阶段产生这一重要因素。我们提出了一种检测框架,通过结合文本检测器、针对每个阶段的验证器、显式的基于规则的風險信号、用户意图与行动一致性分析以及逻辑决策策略来建模这一攻击链。为了支持该框架,我们从提示注入数据集中推导出攻击链标签,在随机划分、时间阶段迁移、条件阶段迁移、跨数据集迁移下评估所提出的框架,并在多个基准上进行消融研究。结果表明,随机训练测试划分在分布偏移下显著高估了鲁棒性,而框架中较后的工具参数阶段比早期阶段更可预测。我们还表明,在类似攻击的无害电子邮件上训练有助于减少误报,同时保持检测真实攻击的能力。在五个二元基准上,我们的框架在严格阈值设置策略下实现了平均F1分数0.406,而未经额外训练的五个预训练检测器中最强的为0.216。这些结果突显了将攻击阶段预测与检索电子邮件中用户请求和指令之间冲突检查相结合的价值。我们的实验还证明了使用具有挑战性的良性示例进行训练以平衡攻击检测和误报的重要性。

英文摘要

Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets, evaluate the proposed framework under random splits, temporal phase transfer, conditional stage transfer, cross-dataset transfer, and conduct ablation studies on multiple benchmarks. Results show that random train test splits substantially overestimate robustness under distribution shift, while later tool argument stages are more predictable than earlier stages in the framework. We also show that training on harmless emails that resemble attacks helps reduce false alarms while preserving the ability to detect real attacks. Across five binary benchmarks, our framework achieves a mean F1 score of 0.406 under the strict threshold setting policy, compared with 0.216 for the strongest of five pretrained detectors evaluated without additional training. These results highlight the value of combining attack stage predictions with checks for conflicts between the user's request and instructions in retrieved emails. Our experiments also demonstrate the importance of training with challenging benign examples to balance attack detection and false alarms.

补充信息

↑