arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniTexture:面向视觉-语言-动作模型的跨任务通用对抗纹理

UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models

Yukun Dai, Mingzhe Dai, Tianshi Wang, Fengling Li, Jingjing Li, Lei Zhu

arXiv 2608.13453首次发表:更新:

AI 中文总结

该研究提出UniTexture跨任务通用对抗纹理攻击,通过单个带纹理3D物体诱导VLA动作预测偏差,在OpenVLA等模型上可显著降低任务成功率,且具备跨套件与跨模型迁移性。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为通用机器人策略,能够遵循多样的语言指令并执行各类操纵任务,但它们对具身智能体的直接控制也使其面临可能引发不安全物理行为的对抗干扰。现有针对机器人策略的攻击通常针对单一任务或指令进行优化,而多任务VLA模型的跨任务漏洞在很大程度上未被探索。我们提出UniTexture,一种跨任务通用对抗纹理攻击,利用单个带纹理的3D物体在多个任务中诱导VLA动作预测的针对性偏差。UniTexture通过可微渲染器将策略动作输出的梯度反向传播至表面纹理参数,利用针对性动作空间目标在任务、指令、状态和视角的分布上联合优化共享纹理,将预测动作导向攻击者定义的目标,无需为每个任务单独优化纹理。我们在OpenVLA和π₀.5上针对多样操纵任务和多种评估设置评估UniTexture,结果显示UniTexture将平均任务成功率从良性条件下的90.0%降至攻击下的48.4%,诱导与目标对齐的动作偏移,且无需重新优化即可进一步展现跨套件和跨模型的迁移性。这些发现共同揭示了多任务VLA模型中可通过单个对抗表面纹理系统利用的共享跨任务漏洞。

英文摘要

Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe physical behaviors. Existing attacks on robotic policies are typically optimized for a single task or instruction, leaving the cross-task vulnerabilities of multitask VLAs largely unexplored. We introduce UniTexture, a cross-task universal adversarial texture attack that uses a single textured 3D object to induce targeted deviations in VLA action predictions across multiple tasks. UniTexture backpropagates gradients from the policy's action outputs to surface texture parameters through a differentiable renderer. It jointly optimizes the shared texture over a distribution of tasks, instructions, states, and viewpoints using a targeted action-space objective, steering predicted actions toward attacker-defined targets without optimizing a separate texture for each task. We evaluate UniTexture on OpenVLA and $π_{0.5}$ across diverse manipulation tasks and multiple evaluation settings. UniTexture reduces the mean task success rate from 90.0% under benign conditions to 48.4% under attack, induces target-aligned action shifts, and further exhibits cross-suite and cross-model transfer without re-optimization. Together, these findings reveal shared cross-task vulnerabilities in multitask VLAs that can be systematically exploited through a single adversarial surface texture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑