arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向元认知少样本间接提示注入:基于结果条件反射的策略抽象

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

Sihan Hou, Xinmeng Hou, Zhijun Zhang, Zehao Wang, Xuhong Ren, Sibo Qin, Kuntharrgyal Khysru, Qing Guo

arXiv 2608.08795首次发表:更新:

发表机构

Nankai University; Nanyang Technological University; Tianjin 712 Mobile Communication Co., Ltd.; Qinghai Minzu University(南开大学; 南洋理工大学; 天津712移动通信有限公司; 青海民族大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SAVOR方法,将攻击适配转向离线策略蒸馏,仅需一次目标智能体查询,在多基准与模型上攻击成功率显著优于现有方法,且记忆可跨防御迁移。

AI 中文摘要

使用工具的大型语言模型(LLM)智能体易受间接提示注入(IPI)攻击,恶意指令嵌入外部观测结果中,以此操纵智能体后续决策与行动。现有多数自适应攻击依赖对目标智能体的反复查询与优化,但现实中攻击者可能仅有一次与未知目标智能体交互的机会。本文提出SAVOR(Strategy Abstraction Via Outcome-Conditioned Reflection,基于结果条件反射的策略抽象),将攻击适配从测试时迭代转向离线策略蒸馏。SAVOR对从不相交训练环境收集的成功与失败轨迹执行结果条件反射,验证上下文条件候选策略,并迭代将其整合为可复用的策略记忆。测试时,冻结的记忆为每个未见过的目标生成单一有效载荷,仅需一次目标智能体查询,无需目标智能体反馈。在两个基准和三个受害模型上,SAVOR在全部6种设置中达到最高平均攻击成功率:相较于最强的现有攻击,其领先2.5至11.8个百分点;在保留攻击者工具的Agent Security Bench上,相较于无策略学习的相同注入通道,其领先23.1个百分点;在本文引入的可执行基准OpenClaw-IPI(保留攻击目标,通过工具交互与执行回执验证攻击)上,其领先28.6个百分点。在一种防御下学习的记忆还可迁移至另一种防御。

英文摘要

Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external observations manipulate subsequent agent decisions and actions. Most existing adaptive attacks rely on repeatedly querying and refining against the target agent, whereas realistic attackers may have only a single opportunity to interact with an unknown target agent. We propose SAVOR (Strategy Abstraction Via Outcome-Conditioned Reflection), which shifts attack adaptation from test-time iteration to offline strategy distillation. SAVOR performs outcome-conditioned reflection over successful and failed trajectories collected from disjoint training environments, validates context-conditioned candidate strategies, and iteratively consolidates them into a reusable strategy memory. At test time, the frozen memory guides the generation of a single payload for each unseen target, requiring only one target-agent query and no target-agent feedback. Across two benchmarks and three victim models, SAVOR attains the highest average attack success rate in all six settings, leading the strongest prior attack by 2.5 to 11.8 points and the same injection channel without strategy learning by 23.1 points on Agent Security Bench, which holds out attacker tools, and 28.6 points on OpenClaw-IPI, an executable benchmark we introduce that holds out attack goals and verifies attacks through tool interactions and execution receipts. A memory learned under one defense also transfers to another.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑