arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分而注入:智能体能否从片段中重构间接提示注入?

Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson

arXiv 2609.36576首次发表:更新:

发表机构

International Computer Science Institute; DSO National Laboratories; UC Riverside; Lawrence Berkeley National Lab(国际计算机科学研究所; 国防科技局国家实验室; 加州大学河滨分校; 劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体系统面临间接提示注入的新威胁,提出自适应长上下文提示注入(AdaLCPI),通过将攻击目标碎片化嵌入检索内容并迭代优化,实现61.4%的攻击成功率,显著高于现有基线,强调安全评估需考虑片段重构场景。

AI 中文摘要

智能体系统现已被广泛用于编排工具和推理长上下文。然而,驱动这些智能体的大型语言模型能力的不断提升,也为间接提示注入创造了新的攻击面。特别是,如果智能体能够从分布在长上下文中的不完整片段重构攻击目标,攻击者可能无需在检索内容中放置完整的恶意指令。在本工作中,我们提出了自适应长上下文提示注入(AdaLCPI),该方法将长上下文碎片化与自适应搜索相结合。AdaLCPI将攻击目标拆分为不完整片段,将其嵌入通过智能体工具检索到的外部内容中,并利用重构提示引导智能体组合这些片段。随后,它通过OpenEvolve,使用分级评分和来自目标智能体的自然语言执行反馈,迭代优化片段和提示。实验表明,AdaLCPI的攻击成功率高于强自适应基线,其宏平均攻击成功率(ASR)达到61.4%,而Trojan Hippo风格基线为32.8%,AgentVigil为30.0%。因此,安全评估应测试当有害目标必须从不完整片段中重构时,智能体是否仍保持稳健性。

英文摘要

Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑