arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对视觉-语言-动作模型的位翻转攻击:动作解码架构决定漏洞

Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability

Yudong Gao, Linghan Chen, Wenhan Wu, Mia Zhou, Jiyao Wang, Kaiyan Ji, Mingyu Guo, Honglong Chen

arXiv 2608.15475首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Adelaide University; Wuhan University; University of North Carolina at Chapel Hill; ETH Zurich; China University of Petroleum (East China)(香港科技大学; 阿德莱德大学; 武汉大学; 北卡罗来纳大学教堂山分校; 苏黎世联邦理工学院; 中国石油大学(华东))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言-动作模型提出位翻转攻击,发现动作解码架构决定漏洞,少量梯度选择的位翻转可大幅降低闭环成功率,权重完整性是具身基础模型的安全边界。

AI 中文摘要

量化视觉-语言-动作(VLA)模型存在权重故障面:Rowhammer式故障可损坏部署的INT8位。我们提出了首个针对VLA的位翻转攻击:少量经梯度选择的翻转可将闭环成功率降至0%,而数百次随机翻转则无危害。在涵盖三类动作头系列的四种模型变体中,有害位集中在少数动作生成层,但经验预算高度依赖动作头类型:直接回归和token策略的翻转次数为1至5次,而评估的流匹配策略则需约100至300次。我们的固定方向流形逃逸损失将\pizero{}的预算从约1000次降至约100次,且匹配的五方向扫描显示该攻击并非特定于全正方向。在直接头模型中,保护3.1%的权重可在K=100时保留60%的成功率,保护5.3%的权重可将开环中断阈值从3次提升至100次。最后,经任务校准的模拟K=100次翻转在真实机器人上的成功率为0/20,而干净模型和全局随机翻转模型的成功率分别为14/20和16/20。因此,权重完整性是具身基础模型的安全边界,代码作为辅助材料提供。

英文摘要

Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We present the first bit-flip attack on a VLA: a few gradient-selected flips reduce closed-loop success to $0\%$, while hundreds of random flips are harmless. Across four model variants spanning three action-head families, damaging bits concentrate in a few action-generating layers, but the empirical budget depends sharply on the head: direct regression and token policies fall in $1$--$5$ flips, whereas the evaluated flow-matching policies require ${\sim}100$--$300$. Our fixed-direction manifold-escape loss cuts \pizero{}'s budget from ${\sim}1000$ to ${\sim}100$ flips, and a matched five-direction sweep shows that the attack is not specific to an all-positive direction. On a direct head, protecting $3.1\%$ of weights preserves $60\%$ success at $K{=}100$, and protecting $5.3\%$ moves the open-loop break threshold from 3 to 100 flips. Finally, task-calibrated emulated $K{=}100$ flips yield $0/20$ real-robot successes, versus $14/20$ clean and $16/20$ global-random. Weight integrity is therefore a security boundary for embodied foundation models. Code is included as ancillary material.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑