arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06862cs.CR

SynChain:诱导计算机使用智能体系统构建自身攻击链

SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains

Fuyao Zhang, Jiaming Zhang, Che Wang, Boyang Chen, Yurong Hao, Xiongtao Sun, Guowei Guan, Blaise Delattre, Yang Cao, Wei Yang Bryan Lim

中文总结 AI 辅助

该研究提出SynChain攻击范式,通过定向监督微调诱导计算机使用智能体构建含恶意内容的良性制品,在多模型多防御设置下实现高攻击成功率,揭示需对智能体跨任务执行轨迹做来源感知推理以保障安全。

中文摘要 AI 辅助

计算机使用智能体(CUAs)已将大型语言模型转化为持久执行系统,能够生成、存储和复用技能、记忆条目等制品。然而,现有安全防御大多将攻击视为外部触发或时间受限的,在应对攻击如何通过智能体自身的持久状态内部传播方面存在关键缺口。我们发现,恶意影响可被隐蔽嵌入自主合成制品的结构冗余中,使其能在内部状态更新中存活并绕过标准审查机制。为形式化这一威胁,我们提出SynChain,一种利用感知持久状态的定向监督微调的自合成攻击范式,用于诱导智能体创建被投毒但外观良性的制品。为系统评估这种传播,我们构建了CUAChain数据集,包含30个良性任务链和3种攻击目标。SynChain能使休眠有效载荷在未来工作流中作为可信上下文无缝重新激活,完全无需新的外部恶意输入。在OpenClaw、Codex和Claude Code上针对4种防御设置开展的大量实验表明,SynChain实现了高攻击成功率且优于适配的基线,证明保护CUAs需要对跨任务执行轨迹进行基于来源的推理。

英文摘要

Computer-use agents~(CUAs) have transformed large language models into persistent execution systems capable of generating, storing, and reusing artifacts like skills and memory entries. However, existing security defenses largely treat attacks as externally triggered or temporally bounded, leaving a critical gap in addressing how compromise can propagate internally through an agent's own persistent state. We reveal that malicious influence can be covertly embedded into the structural redundancies of autonomously synthesized artifacts, allowing it to survive internal state updates and bypass standard vetting mechanisms. To formalize this threat, we introduce SynChain, a self-synthesized attack paradigm utilizing persistence-aware directed supervised fine-tuning to induce agents to create poisoned yet benign-looking artifacts. To systematically evaluate this propagation, we construct CUAChain, a dataset comprising 30 benign task chains and three attack objectives. SynChain enables dormant payloads to seamlessly reactivate in future workflows as trusted context, operating entirely without new malicious exogenous inputs. Extensive experiments on OpenClaw, Codex, and Claude Code under four defense settings demonstrate that SynChain achieves high attack success and outperforms adapted baselines, proving that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.

↑