arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

灯光、相机、故障:当光照鲁棒性使视觉语言动作模型对颜色视而不见时

Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color

Marino Watanabe, Takami Sato, Kentaro Yoshioka

arXiv 2607.14698首次发表:更新:

发表机构

Keio University(庆应义塾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究VLA模型在现实环境中对光照干扰的脆弱性,提出FLARE攻击框架和ChromaGuard对抗训练方法,通过实验验证ChromaGuard能有效提升模型在颜色相关任务中的成功率,保障模型在复杂环境下的性能。

AI 中文摘要

视觉语言动作(VLA)模型已成为通用机器人操纵的强大范例,但向现实环境过渡时易受微小环境干扰影响。我们提出FLARE,一种优化的物理聚光灯攻击框架,通过有针对性的光照利用这些漏洞,能使基线任务成功率降至零且无需访问模型内部。对抗训练是标准对策,但存在防御陷阱,朴素数据增强会使VLA模型将颜色视为噪声而丢弃,导致视觉感知退化。我们通过诊断灰度评估揭示了这种退化,提出了色度保护对抗训练方法ChromaGuard。在物理6自由度机器人平台上,ChromaGuard在良性和受攻击的颜色相关任务中的成功率分别达到97.5%和92.5%。

英文摘要

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general-purpose robot manipulation; however, their transition to real-world environments reveals vulnerabilities to minor environmental perturbations. We propose FLARE, an optimized physical spotlight attack framework that exploits these vulnerabilities via targeted illuminations, dropping baseline task success rates to zero without any access to model internals. While adversarial training is the standard countermeasure, we identify a critical and previously underestimated defensive pitfall: naive data augmentations incorrectly condition VLA models to discard color as noise, collapsing their visual perception into a purely shape-biased processor. We expose this degradation through a diagnostic grayscale evaluation, in which the defended model maintains high success rates on grayscale inputs, while its success rate on benign, color-dependent real-world tasks drops to at most 47.5%, well below the undefended baseline. To address this, we propose ChromaGuard, a chroma-preserving adversarial training method. On a physical 6-DoF robotic platform, we demonstrate that ChromaGuard achieves 97.5% and 92.5% success rates in benign and attacked color-dependent tasks, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑