arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐形墨水威胁:计算机使用智能体中合法任务背后的对抗目标

Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

Jia-Chen Zhang, Ze-Yu Zhang, Kai-Wei Zhang

arXiv 2608.02018首次发表:更新:

AI 中文总结

该研究针对计算机使用智能体的隐形墨水威胁,构建II-Bench与HITLCUA框架评估后发现,低危害注入可绕过防御,暴露出现有CUAs的安全风险。

AI 中文摘要

计算机使用智能体(Computer-use agents, CUAs)可赋能大型语言模型自主操作系统及网页,如今正日益易受间接提示注入攻击。一种被广泛采用的防御机制是人类在环(human-in-the-loop)范式,即智能体在执行敏感操作前会暂停以等待明确的用户确认。尽管该机制能有效抵御危害明显的攻击,但对我们所称的隐形墨水威胁(Invisible Ink Threats)几乎无防护作用:这类注入的目标危害较低,例如为代码仓库加星或安装软件包,其行为与合法任务执行无法区分,因此能规避模型安全机制与人类监督。为系统探究这一盲区,我们推出II-Bench,一组看似无害的对抗任务集合。II-Bench包含444个针对机密性与完整性攻击的示例,覆盖三个平台,分为三类攻击:页面导航与交互、敏感信息 exfiltration(泄露)、代码下载与执行。每类攻击均以自然语言和代码形式呈现,且包含两种指令明确程度。此外,我们构建HITLCUA,一个综合对抗测试框架,将真实虚拟机操作系统环境与基于Docker的隔离网页平台整合,并通过允许CUAs在可疑操作前咨询API模拟用户来模拟人类参与。对主流CUAs的广泛评估显示,低危害注入常能绕过智能体防御机制与模拟用户审查,暴露出当前CUAs中严重且此前未被充分探索的安全风险。

英文摘要

Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection attacks. A widely adopted defense is the human-in-the-loop paradigm, in which the agent pauses for explicit user confirmation before executing sensitive operations. While effective against conspicuously high-harm attacks, this defense offers little protection against what we term Invisible Ink Threats: low-harm injected goals, such as starring a repository or installing a package, that are behaviorally indistinguishable from legitimate task execution and thus evade both model safety mechanisms and human oversight. To systematically investigate this blind spot, we present II-Bench, a collection of seemingly harmless adversarial tasks. II-Bench comprises 444 examples targeting confidentiality and integrity attacks across three platforms, spanning three attack categories: page navigation and interaction, sensitive information exfiltration, and code download and execution. Each category is instantiated in both natural language and code forms under two levels of instruction specificity. Furthermore, we construct HITLCUA, a comprehensive adversarial testing framework that integrates a real virtual machine operating system environment with isolated Docker-based web platforms, and simulates human participation by allowing CUAs to consult an API-simulated user before proceeding with suspicious operations. Extensive evaluations of leading CUAs reveal that low-harm injections frequently bypass both agent defenses and simulated user review, exposing severe and previously underexplored security risks in current CUAs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑