arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09212cs.CRcs.AIcs.CY

AgentHijack:多模态计算机使用代理的视觉补丁攻击

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

Zhihao Liu, Hongyu Sun, Zhiyuan Fu, Xiaonan Duan, Jice Wang, Shangru Zhao, Weizhi Meng, Wuxin Yang, Yangfan Zhou, Yuqing Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AgentHijack框架,通过视觉补丁攻击多模态计算机使用代理,实验显示可诱导代理执行恶意命令,造成真实环境风险。

中文摘要 AI 辅助

本文提出了一种针对计算机使用代理(CUA)的图像触发命令注入的端到端评估框架。其目标是测试局部视觉补丁能否在截图输入、VLM生成、动作解析和环境执行的完整链条中诱发可验证的环境后果。我们在作者控制的GitHub Pages页面和本地部署的CSDN克隆上训练并部署补丁,并在五个开源或公开可用的GUI代理或视觉语言模型(VLM)后端上于真实环境中进行评估。我们的实验汇总了600个实例级在线案例,其中T-ASR、TAPR和E2E-ASR分别达到84.5%、47.0%和20.3%。轨迹分析进一步表明,在一些成功案例中,代理首先执行恶意终端命令,然后继续原始良性任务。这些结果表明,优化的局部视觉信号不仅能影响VLM输出,还能通过开放CUA的执行管道传播,并造成真实的环境风险。

英文摘要

This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input, VLM generation, action parsing, and environment execution. We train and deploy patches on author-controlled GitHub Pages pages and a locally deployed CSDN clone, and evaluate them in real environments across five open-source or publicly available GUI-agent or vision-language-model (VLM) backends. Our experiment aggregates 600 instance-level online cases, with T-ASR, TAPR, and E2E-ASR reaching 84.5%, 47.0%, and 20.3%, respectively. Trajectory analysis further shows that in some successful cases the agent first executes a malicious terminal command and then continues the original benign task. These results indicate that optimized local visual signals can affect not only VLM outputs but also propagate through the execution pipeline of open CUAs and create real environmental risk.

发表机构

  • Hainan University(海南大学)
  • University of Chinese Academy of Sciences National Computer Network Intrusion Protection Center(中国科学院大学国家计算机网络入侵防护中心)
  • Lancaster University(兰卡斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑