发表机构
National University of Singapore; The Hong Kong Polytechnic University(新加坡国立大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对世界-动作模型(WAMs)提出BadWAM框架,用于建模和评估世界-动作漂移攻击,包括仅动作攻击和保持想象攻击,通过不同标准刻画攻击面,评估结果显示能大幅降低任务成功率,揭示WAM漏洞。
AI 中文摘要
世界-动作模型(WAMs)正成为具身控制的一个有前景的基础:它们学习将动作生成与未来世界预测相结合的表示,而非仅预测动作。这种耦合常被视为鲁棒性、可解释性和安全性的来源。本文表明该假设是脆弱的。我们引入BadWAM,一个用于建模和评估世界-动作漂移攻击的统一框架,这是一类新的针对WAM的对抗攻击,利用小的视觉扰动打破WAM想象与执行之间的对齐。BadWAM根据攻击强度和隐蔽性这两个自然标准来刻画这种攻击面。当对手优先考虑破坏时,BadWAM实例化仅动作对抗攻击,直接驱使模型走向导致任务失败的动作。当对手还优先考虑隐蔽性时,BadWAM实例化保持想象的对抗攻击,试图在使模型预测的未来接近其纯净想象的同时引发有害的动作转变。我们在不同的WAM变体上评估BadWAM。结果表明,我们的攻击在闭环执行下大幅降低了任务成功率。例如,我们的仅动作攻击将模型性能从96.5%的成功率降至43.1%。我们的保持想象攻击结果进一步揭示了WAM特有的漏洞:适度的未来保持正则化可以保持强大的攻击性能,同时减少未来想象漂移。
英文摘要
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5\% to 43.1\% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.