arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉-语言-动作模型对单步观测扰动是否具有鲁棒性?

Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?

Shojiro Yamabe, Jun Sakuma

arXiv 2609.32550首次发表:更新:

AI 中文总结

本研究针对VLA模型在单步观测扰动下的鲁棒性不足问题,提出CARE方法,通过动态调整动作块执行长度,以低开销提升鲁棒性并保持清洁性能。

AI 中文摘要

理解视觉-语言-动作(VLA)模型的安全风险对于其在物理世界中的部署至关重要。现有的安全研究主要考虑持续扰动,即在整个回合中连续施加于观测的扰动。然而,瞬时观测损坏(即观测仅在回合内短暂地受到严重扰动)仍是一个未被充分探索的安全威胁。为弥补这一空白,本研究探讨了在每回合单个时间步施加的单步扰动的鲁棒性。我们的实验表明,这些扰动显著降低了VLA的性能,且其影响取决于动作块执行长度。基于此,我们提出了CARE,该方法根据与先前预测的动作块的一致性动态选择执行长度。CARE通过仅在扰动下选择较短的执行长度,以较低的计算开销提高鲁棒性,同时保持清洁性能。

英文摘要

Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which observations are severely perturbed only briefly within an episode, remains an underexplored safety threat. To address this gap, this work investigates robustness to one-step perturbations applied at a single time step per episode. Our experiments reveal that these perturbations substantially degrade VLA performance and that their impact depends on the action chunk execution length. Based on them, we propose CARE, which dynamically selects the execution length based on consistency with the previously predicted action chunk. CARE improves robustness with low computational overhead while preserving clean performance by selecting shorter execution lengths only under perturbations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑