arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAPDreamer:面向世界动作模型的可迁移对抗补丁

TAPDreamer: Transferable Adversarial Patches for World Action Models

Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng, Yichen Feng, Yaorui Ding, Yuetai Li, Zhen Xiang, Bhaskar Ramasubramanian, Basel Alomair, Luyao Niu, Radha Poovendran

arXiv 2610.06814首次发表:更新:

发表机构

University of Washington; University of Georgia; Western Washington University; King Abdulaziz City for Science and Technology; HUMAIN(华盛顿大学; 佐治亚大学; 西华盛顿大学; 阿卜杜勒阿齐兹国王科技城; HUMAIN)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TAPDreamer攻击,仅用公共编码器构建可迁移的固定局部补丁,无需目标策略查询,通过最大化全局L1距离使FastWAM成功率降至0%,揭示防御需保护共享视觉编码器。

AI 中文摘要

世界模型学习预测其环境将如何演变,使其成为通用机器人控制的重要基础。然而,世界动作模型依赖于摄像头输入,对这些输入的操纵可能会破坏跨任务和动作策略使用的视觉表示。针对这些模型的现有攻击针对受害者的动作或预测的未来进行优化,因此需要访问目标模型的输出。在本文中,我们提出了一种针对世界动作模型的攻击方法TAPDreamer,该方法仅使用公共编码器来构建一个固定的局部扰动,该扰动可跨任务和动作架构迁移。TAPDreamer不需要目标策略查询。我们的关键洞察是,补丁引起的注意力权重和值向量变化之间的相互作用,会将几乎相同的表示偏移传播到远超补丁覆盖范围的区域,并且这种偏移在不同任务观测中保持稳定。基于这一洞察,TAPDreamer使用来自一个源任务的六帧图像来最大化干净与补丁编码器表示之间的全局L1距离。在闭环评估中,每个基准使用一个冻结补丁,覆盖输入约6.5%的区域,将FastWAM在40个LIBERO任务上的成功率从97.7%降至0.0%,在50个RoboTwin任务上的成功率从90.8%降至0.0%;匹配的随机补丁分别保留了81.5%和79.2%的成功率。相同的补丁在两种DreamWAM配置上将成功率降至2.1%和0.8%,在Motus上将成功率降至10.0%。这些结果表明,仅保护下游动作生成是不够的:世界动作模型的防御必须同时保护共享视觉编码器免受持久局部扰动的影响。

英文摘要

World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.86% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 1.45% and 1.00% on two DreamWAM configurations and to 10.60% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.

CommentsProject Page: https://tapdreamer.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑