发表机构
University of Utah(犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉定位网页代理,提出WebMirage框架,通过局部视觉扰动实现端到端劫持,平均攻击成功率91.9%,显著优于基线,且能抵御多种防御。
AI 中文摘要
基于大型视觉语言模型的现代网页代理处理网页、选择相关UI元素,并将模型输出转化为浏览器操作。现有的视觉红队方法使用对抗性视觉内容来操纵这一过程。然而,这些方法主要针对模型推理,并未明确考虑结构化输入处理或操作后处理。因此,模型层面的成功并不能确保对浏览器执行的控制,也无法可靠地表征端到端代理的鲁棒性。为解决这一空白,我们将视觉定位网页代理的红队测试形式化为一个端到端的“定位到执行”问题,并引入WebMirage框架,该框架制作局部视觉扰动,使代理选择攻击者控制的内容,并在不同网页渲染下执行相应的浏览器操作。它使用角色槽抽象和网页重组来捕捉网页元素间的竞争,并利用数据流分析使优化与操作后处理对齐。我们在四种代理配置和六种VLM骨干网络上,对涵盖13个公共网站和一个沙箱基准的2250个任务进行了WebMirage评估。WebMirage实现了91.9%的平均攻击成功率,而最强基线仅为17.4%,并且对三种代理级防御仍然有效。
英文摘要
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.
Comments20 pages, 8 figures, 6 tables. Code: https://github.com/MoonTea0416/WebMirage