arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ZoneClaw:通过在OpenClaw式计算机使用智能体中建立内存分区来缓解持久性内存攻击

ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents

Haokai Ma, Chieh Lin, Yupeng Qiu, Ee-Chien Chang

arXiv 2610.00450首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ZoneClaw通过分层信任区分离持久性与权威,阻止外部声明获得行动指导权,将攻击成功率从372/480降至6/480,同时保持高实用性。

AI 中文摘要

计算机使用智能体日益通过持久性工作区内存作为长期运行的助手运行,OpenClaw式CUA将其实现为自动重新加载的文件,这些文件以相同的权限级别保存用户指令、系统摘要和外部声明。在此,记住一个声明即赋予对后续行为的权威。这引发了一种持久性内存攻击,在这种攻击中,仅控制良性外观外部内容的攻击者诱导CUA在合法任务期间记录有利于攻击者的声明,而这些声明随后支配攻击者从未触及的良性任务。攻击链将“恶意上下文→恶意响应”扩展为“恶意上下文→内存注入→恶意执行”,使其成为一种跨环境威胁。现有防御研究要么在内容进入内存之前干预,要么在其随后诱导的动作处干预,而非关注存储内容是否可能指导动作。我们提出ZoneClaw,它通过将扁平工作区内存替换为带有明确权威级别的分层信任区来将持久性与权威分离。外部声明持久存在于低信任区,并且仅通过跨越明确的权威边界才获得行动指导权威,在此边界处,提升会与攻击者无法直接写入的区域进行交叉检查。非对称权限的角色特定进程强制执行此边界,确保没有进程既摄取外部内容又向外行动。在四种攻击场景、两种注入设置和四种骨干模型下,ZoneClaw将ASR从372/480降至6/480,同时在458/480次试验中保持实用性,并且对某些具有防御意识的攻击者仍然有效。攻击者的声明仍然持久存在于低信任内存中,但很少跨越权威边界,表明ZoneClaw是扣留权威而非拒绝从环境中学习。我们的代码可在以下网址获取:this https URL。

英文摘要

Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions, system summaries, and external claims at the same privilege level. Here, remembering a claim confers authority over later behavior. This enables a persistent memory attack, in which an attacker who controls only benign-looking external content induces the CUA to record attacker-favored claims during a legitimate task, and those claims later govern benign tasks the attacker never touches. The attack chain extends "malicious context -> malicious response" into "malicious context -> memory injection -> malicious execution", making this a cross-environment threat. Existing defenses studied intervene either before content enters memory or at the action it later induces, not whether stored content may guide action. We propose ZoneClaw, which separates persistence from authority by replacing flat workspace memory with hierarchical trust zones carrying explicit authority levels. External claims persist in a low-trust zone and acquire action-guiding authority only by crossing an explicit authority boundary, at which promotion is cross-checked against zones the attacker cannot directly write. Role-specific processes of asymmetric privilege enforce this boundary, ensuring that no process both ingests external content and acts outward. Across four attack scenarios, two injection settings, and four backbones, ZoneClaw drives ASR from 372/480 to 6/480 while retaining utility in 458/480 trials, and remains effective against some defense-aware attackers. Attacker claims still persist in low-trust memory yet rarely cross the authority boundary, showing that ZoneClaw withholds authority rather than refusing to learn from the environment. Our code is available at: https://github.com/euph00/ZoneClaw-code.

Comments34 pages, 9 figures; Under Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑