arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WeClawArena:面向以人为中心的智能体网络中跨用户智能体协作与安全的可审计沙箱及基准

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang

arXiv 2608.03499首次发表:更新:

AI 中文总结

WeClawArena是面向以人为中心的智能体网络的可审计沙箱与基准,含124个基础任务及620个场景变体,可评估跨用户智能体协作的效用与攻击成功率,支持安全审计与诊断。

AI 中文摘要

近期,持久化个人智能体框架的进展使得以人为中心的智能体网络成为现实部署目标:每个用户可由代表其行事、维护状态并通过社交与任务关系与其他智能体通信的AI智能体提供服务。在这些网络中,日常工具使用转变为跨个人工作空间的多拥有者智能体协作,其中文件、记录、工具和策略无法被其他拥有者直接查看。现有智能体基准研究了工具使用与协作,但未提供用于可验证跨用户智能体协作的端到端沙箱,该沙箱需具备现实用户数字工作空间,也未测试有害行为如何在以人为中心的智能体网络中传播。本文提出WeClawArena,这是一个面向跨个人工作空间的多拥有者智能体协作的可审计基准与运行时沙箱。WeClawArena针对协作工具使用任务,其中个人工作空间兼具操作工具与个人约束的作用。该基准包含6个跨用户任务领域的124个基础任务,并将其扩展为620个场景变体,每个基础任务对应1个良性对照组和4个攻击向量变体。沙箱会记录对等消息、工具调用、资源操作、受控决策以及最终工作空间状态。WeClawArena分别报告效用和攻击成功率,并通过受限运行时证据审计攻击成功情况,支持诊断任务故障、隐私泄露、中毒证据及无效权限路径。

英文摘要

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.

Comments31 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑