arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证而非采样:视觉-语言模型与视觉-语言-动作模型的区域级鲁棒性

Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models

Bogdan Aron, Christopher Brix, Benedikt Brückner, Yanghao Zhang, Panagiotis Kouvaros, Alessio Lomuscio

arXiv 2609.22293首次发表:更新:

发表机构

Safe Intelligence(Safe Intelligence)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出H$^2$V-M验证框架,对多个VLM和VLA在连续图像扰动区域进行鲁棒性验证,发现鲁棒性主要取决于扰动类型,且VLM对旋转更鲁棒,VLA对微小扰动敏感。

AI 中文摘要

视觉-语言模型(VLMs)和视觉-语言-动作模型(VLAs)正越来越多地部署于现实世界应用中。在这些应用中,对记录到的相机图像进行微小扰动可能显著改变决策。然而,现有针对这些模型的基准测试仅对扰动进行采样,这无法保证在未测试区域内不存在失败。我们首次对六个VLM(来自Gemma、InternVL、LLaVA和Qwen系列)和五个VLA(来自GR00T、OpenVLA和π系列)在完整连续的光度和几何图像扰动区域(包括亮度偏移、相机旋转及其组合)上进行了鲁棒性验证。为此,我们基于验证框架H$^2$V,并引入H$^2$V-M,一种考虑边界的收敛规则,使得在32B参数规模下的验证变得可行。我们证明,H$^2$V-M在模型查询次数上比H$^2$V提升一个数量级,并且比随机采样更快地找到反例,同时提供可靠性保证。我们的VLM和VLA鲁棒性验证表明,鲁棒性主要依赖于扰动类型而非模型本身,并且VLM对较大相机旋转的鲁棒性优于VLA。对于VLA,即使像±1°这样小的扰动也可能在许多情况下改变所指令的动作。我们还表明,鲁棒性更依赖于模型系列而非模型规模。

英文摘要

Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the absence of a failure in the untested region. We present the first robustness validation of six VLMs (drawn from the Gemma, InternVL, LLaVA, and Qwen families) and five VLAs (drawn from the GR00T, OpenVLA, and $π$ families) over entire continuous regions of photometric and geometric image perturbation: brightness shifts, camera rotations, and their composition. To this end, we build on the validation framework H$^2$V and introduce H$^2$V-M, a margin-aware convergence rule that makes validation affordable at the 32B parameter scale. We demonstrate that H$^2$V-M outperforms H$^2$V by an order of magnitude in model queries and that it finds counterexamples faster than random sampling while providing soundness guarantees. Our VLM and VLA robustness validation shows that robustness is mostly dependent on the perturbation type, rather than the model, and that VLMs are more robust to large camera rotations than VLAs. For VLAs, even perturbations as small as $\pm1^\circ$ can change the commanded action in many cases. We also show that robustness depends more on model family than on model size.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑