双视图:防止个人人工智能代理中的间接提示注入
DualView: Preventing Indirect Prompt Injection in Personal AI Agents
- Seoul National University(首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究个人人工智能代理面临的间接提示注入攻击问题,提出双视图方法,通过扩展不可信数据跟踪到用户环境,部署为插件,有效阻止攻击并保持实用性。
AI中文摘要:
在用户本地机器上运行的个人人工智能代理,如OpenClaw,会面临间接提示注入(IPI)攻击。先前的双大语言模型防御通过用代理可引用但不可读的符号替换不可信数据来阻止IPI,但存在存储IPI问题。本文提出DualView,将不可信数据跟踪扩展到用户环境,部署为OpenClaw插件,能阻止各种IPI攻击。
英文摘要:
Personal AI agents that run on the user's local machine automate daily tasks including web search, email, and file management. Their access to computer resources, including the network, file system, and shell, exposes them to indirect prompt injection (IPI) attacks. Prior Dual LLM defenses block IPI by replacing untrusted data with symbols that the agent can reference but not read. However, they track untrusted data only inside the agent's context, so when the agent saves and later rereads untrusted data, that data, possibly an attacker's prompt, can return as trusted data rather than as a symbol, which we call stored IPI. Operating on the user's real environment is what makes agents like OpenClaw practical, and is exactly why a defense that ignores it is incomplete. Preserving symbols in such an environment is hard, because humans and programs need original data. We present DualView, which extends untrusted data tracking from the agent's context to the user's environment, including the file system, shell, network, and other agents, by giving each channel two views. In AgentView, the agent sees untrusted data as symbols even after writing it out and reading it back, blocking stored IPI, while HumanView preserves original data for humans and tools. DualView routes each tool call to the right view and synchronizes data across the two views. DualView deploys as an OpenClaw plugin using only tool hooks, without changing the agent's tool-call logic or tool implementations. DualView deterministically prevents instructions in untrusted data from directly steering the agent's tool calls; this guarantee does not depend on recognizing the evaluated attack templates. In our evaluation on an IPI benchmark and PinchBench, DualView blocked every tested IPI attack, including stored IPI. On PinchBench, its utility drop was within 1.8 to 6.4 points.